%my 5 2.1 X. POLICY RESEARCH WORKING PAPER 2672 Do Workfare Participants A lot can be learned about the impact of an antipoverty Recover Quickly from program by studying income Retrenchment? replacernent for those observed to leave the program after its Martin Ravallion retrenchment. A Bank- Emanuela Galasso supported workfare program Teodoro Lazo in Argentina is found to have Ernesto Philipp had a sizable impact on participants' incomes. The World Bank Development Research Group Poverty September 2001 | POLICY RESEARCH WORKING PAPER 2672 Summary findings What happens to participants in a workfare program-a up one quarter of the gross workfare wage within six program that imposes work requirements on welfare months. This rises to half in 12 months. The estimates recipients-when that program is cut? Ravallion, are unbiased in the presence of time-invariant errors Galasso, Lazo, and Philipp compare the incomes of from mismatching in the selection of the comparison workfare participants in Argentina to those of group. Fully removing selection bias would probably nonparticipants and past participants after a severe yield even lower income replacement. Test results based contraction in aggregate outlays on the program. The on a second follow-up survey suggest that valid authors find evidence of partial income replacement, inferences can be drawn about program impacts from the such that those who left the program were able to make authorsi measures of income replacement. This paper-a product of the Poverty Team, Development Research Group-is part of a larger effort in the group to assess the impact of Bank-supported antipoverty programs. The study was funded by the Bank's Research Support Budget under the research project "Policies for Poor Areas" (RPO 681-39). Copies of this paper are available free from the World Bank, 1818 H Street NW, Washington, DC 20433. Please contact Catalina Cunanan, room MC3-542, telephone 202-473-2301, fax 202-522-1151, email address mravallion@worldbank.org. Policy Research Working Papers are also posted on the Web at http://econ.worldbank.org. The authors may be contacted at mravallion@worldbank.org or egalasso@worldbank.org. September 2001. (35 pages) The Policy Research Working Paper Series disseminates the findings of work in progress to encourage the exchange of ideas about development issues. An objective ofthe series is to get the findings out quickly, even if the presentations are less than fully polished. The papers carry the names of the authors and should be cited accordingly. The findings, interpretations, and conclusions expressed in this paper are entirely those of the authors. They do not necessarily represent the view of the World Bank, its Executive Directors, or the countries they represent. Produced by the Policy Research Dissemination Center Do workfare participants recover quickly from retrenchment? Martin Ravallion, Emanuela Galasso, Teodoro Lazo, Ernesto Philipp' Development Research Group Trabajar Project Office World Bank Ministry of Labor Government ofArgentina Keywords: workfare; propensity-score matching; double-difference; Argentina JEL classifications: H43, I38 I The work reported in this paper is part of the ex-post evaluation of the World Bank's Social Protection III Project in Argentina. The authors' thanks go to staff of the Trabajar project office in the Ministry of Labor, Government of Argentina, who have helped in countless ways, and to the Bank's Manager for the project, Polly Jones, for her continuing support of the evaluation effort, and many useful discussions. We also benefited from discussions with Jyotsna Jalan and the comments of Guilermo Perry and seminar participants at the Ministry of Labor. Support from the Evaluation Thematic group of the World Bank's Poverty Reduction and Economic Management Network is gratefully acknowledged. These are the views of the authors, and need not reflect those of the Government of Argentina or the World Bank. Correspondence: Martin Ravallion, World Bank, 1818 H Street NW, Washington DC, 20433 USA; mravallion@worldbank.org. 1. Introduction The welfare outcomes of cutting a workfare program-which imposes work requirements on welfare recipients-will depend in part on labcir market conditions facing the participants. High income replacement after retrenchment might suggest that unemployment is not a serious poverty problem. But even when there is high unemployment, there are other ways that retrenched workers might recover the lost income. Possibly the work experience on the program will help them find work, including self-employment. Or possibly private transfers will help make up for the loss of public support. Tracking ex-participants after their retrenchment and measuring their income replacement may thus provide important clues to understanding the true impact of a workfare program. This paper tires to learn about the impact of a workfare program by studying income replacement for those observed to leave the program after its contraction. The analytic problem we face is the usual one in causal studies of missing data on the counter-factual. It is well recognized that "single difference" comparisons of income levels between participants and non- participants can be highly misleading given the existence of (observable and unobservable) heterogeneity in characteristics that jointly influence participation and incomes in the absence of the program. Simulations and comparisons with actual experiments have suggested that careful matching in terms of observable covariates can greatly reduce the bias in observational studies.2 Amongst the various matching methods available, Propensity Score Matching (PSM) has attracted recent interest given its theoretical properties, notably that exact matching by this 2 For evidence based on simulations see Rubin (1979) and Rubin and Thomas (2000). For evidence based on an actual evaluation see Dehejia and Wahba (1999) who find that single-difference matching based on propensity scores gives a good approximation to results of a randomized evaluation of a US training program - much better than the non-experimental methods that Lalonde (1986) assessed for the same program. However, Smith and Todd (2000) question the robustness of Dehejia and Wahba's findings to model specification. 2 method is the observational equivalent of randomization (Rosenbaum and Rubin, 1983, 1985). PSM gives unbiased estimates if (inter alia) the conditional independence ("strong ignorability") assumption holds, whereby pre-intervention outcomes are independent of participation given the variables used for matching (Rosenbaum and Rubin, 1983). Conditional dependence will leave a bias, which will depend on the amount of relevant data available for matching. Another approach in the literature is the popular double difference (DD) estimate, obtained by comparing treatment and comparison groups in terms of outcome changes over time relative to a pre-intervention baseline. DD allows for conditional dependence arising from additive time-invariant latent heterogeneity. Since PSM optimally balances observed covariates between the treatment and comparison groups, it is the obvious method for selecting the comparison group in double-difference studies. The results of Heckman, Ichimura and Todd (1997), Heckman et al., (1998), Heckman and Smith (1998) and Smith and Todd (2000) suggest that a hybrid method, combining PSM for selecting the comparison group with DD to eliminate time-invariant errors, can greatly reduce (but not eliminate) the bias found in other evaluation methods, including single-difference matching. DD estimators have their limitations. In some circumstances it is implausible that the selection-bias is time invariant. For example, there is a potential bias in DD estimators when the changes over time are a function of initial conditions that also influence program placement.4 There is also the well-known bias for inferring long-term impacts that can arise when there is a 3 For recent evidence based on simulations see Rubin and Thomas (2000). For evidence based on actual evaluation see Dehejia and Wahba (1999) who find that single-difference PSM given a good approximation to results of a randomized evaluation of a US training program - much better than the non-experimental methods that Lalonde (1986) assessed for the same program. Smith and Todd (2000) question the robustness of Dehejia and Wahba's findings to model specification. 4 Jalan and Ravallion (1998) show that this can seriously bias evaluations of poor-area development programs that are targeted on the basis of initial geographic characteristics that also influence the growth process. 3 pre-program earnings dip (known as "Ashenfelter's dip" following Ashenfelter, 1978). In assessing short-term impact - a common concern of safety-net interventions -one would not normally want to ignore this dip, though it r emains relevant to assessing the time profile of gains from the safety net. What if one does not have a pre-intervention baseline? This is common for safety-net interventions, such as workfare programs, that have to be set up quickly, in response to a macroeconomic or agro-climatic crisis. There is no time to do a baseline survey of (probable) participants and non-participants. Nor is ranidomization usually feasible in such settings. Suppose instead that we follow up samples of participants and non-participants over time, post- intervention, and that some participants become non-participants. What can we then learn about the program's impacts? The approach we propose here is to examine what happens to participants' incomes when they leave a workfare program, and to compare this with the incomes of continuing participants, after netting out economy-wide changes, as revealed by a matched comparison group of non- participants. While this approach is feasible without a baseline survey, it brings its own problems. Firstly, while differencing over time can eliminate bias due to latent (time-invariant) matching errors, there remains a potential bias due to any selective retrenchment from the program based on unobservables. We argue that the direction of bias can be determined under plausible assumptions. Secondly, while we are not concerned with any pre-program "Ashenfelter's dip," there may well be a post-program version of the same phenomenon, namely when earnings drop sharply at retrenchment, but then recover. As in the pre-program dip, this need not be a source of bias in assessing the impact of a safety-net intervention (to the extent that the pre-program dip 4 entails a welfare change); nonetheless, the post-program dip is clearly of interest in assessing the dynamics of recovery from retrenchment. To help address this issue we follow up initial participants over multiple survey rounds. We are also interested in seeing whether this type of follow-up study of participants can identify the gains to current participants from a program - the classic "treatment effect on the treated" as it is called in the evaluation literature. There are concerns about selection bias, and there is the problem that past participation may bring current gains to those who leave the program. Assuming these lagged gains are positive, the net loss from leaving the program will be less than the gain from participation relative to the counter-factual of never participating. We derive a test for the joint conditions needed to identify the mean gains to participants from this type of study, also exploiting further follow-up surveys of past participants. We study Argentina's Trabajar Program. This government program aims to provide work to poor unemployed workers on approved sub-projects of direct value to poor communities. The sub-projects cannot last more than six months, though a worker is not prevented from joining a new project, if available. In earlier research on the same program, Jalan and Ravallion (1999) estimated the counter-factual income of current participants if they had not participated using the mean income of a comparison group of non-participants, obtained by PSM. For the purpose of the present study, we designed a survey of a random sample of current participants, and returned to the same households six months later, and then 12 months later. In addition to natural rotation, there was a very sharp contraction in the program's aggregate outlays after the first survey. The following section describes the program and the data for its evaluation. Section 3 describes our evaluation method in theoretical terms. Section 4 presents our results, while some conclusions can be found in Section 5. 5 2. The program and data In response to a sharp increase in the measured unemployment rate, the Government of Argentina greatly expanded and redesigned its Trabajar Program in May 1997, with financial and technical support from the World Bank. The Trabajar Program aims to provide short-term work at relatively low wages on socially useful projects in poor areas. The projects are proposed by local (governmental and non-governmental) organizations with priority given to proposals that are likely to benefit poor areas, according to ex-ante assessments. Workers cannot join the program unless they aie recruited to an approved project. The projects last a maximum of six months, but a worker is not prevented from switching to a new project on the same basis. The wage rate was initially set at a rnaximum of $200 per month, which was cut to $160 in 1999 at the time of an overall contraction in outlays. (Undercutting of the wage rate is allowed, but it is uncommon.) The wage rate was chosen to be low enough to assure good targeting performance, and to help assure workers would take up regular work when it became available. By way of comparison, the average monthly earnings for workers in the poorest 10% of households (ranked by total income per person) in Greater Buenos Aires (GBA) in May 1996 was $263 (calculated from the Permanent Household Survey, discussed further below). (As expected, the poorest decile also received the lowest average wage, and average wages rose monotonically with household income per person.) The data collection for this study began with a survey in May/June 1999 of Trabajar participants in the main urban areas of three provinces - Chaco (Gran Resistencia), Mendoza (Gran Mendoza) and Tucuman (Gran Tucaman - Tafi Viejo). These provinces were chosen as representing the range of labor markets found in Argentina. The families of 1500 randomly chosen Trabajar workers were interviewed, spread evenly between the three provinces. The 6 sampled beneficiary households were a simple random sample from the list of all beneficiaries at the time. The households of participants were the units for interviewing. The survey of participants was chosen to coincide with the twice-yearly Permanent Household Survey (PHS). This is an urban survey focusing on employmnent and incomes, though it also includes questions on education and demographics. We calculate individual income from questions on income from work (wages, bonuses, self-employment income, Trabajar earnings) and from non-labor sources (pension, rents, dividends, fellowships, food coupons, private transfers). All provincial capitals or other urban centers with at least 100,000 inhabitants are included in the PHS.5 The survey is conducted twice a year, around May and October. The PHS sample size is set to achieve (with 95% confidence) an error of 1% in the unemployment rate within each urban conglomerate. In large conglomerates, a random sample of geographic units is chosen, within which a fixed number of households is sampled. In smaller conglomerates, a one- stage random sample is used. The PHS sample includes 27,000 households. The PHS is our source of the comparison group for initial participants, to be selected by propensity-score matching, as described in more detail in the next section. For program participants, the same interview questionnaire was used as for the PHS, with PHS interviewers. This avoids the matching bias that can arise when the surveys of participants and non- participants are not comparable (Heckman, Ichimura and Todd, 1997; Heckman et al., 1998). Extra questions were added for the survey of Trabajar participants. Moreover, miss-matching can be reduced by selecting the comparison group separately from each geographic area, to make sure that the individuals belong to the same local labor markets.6 5 An exception is Viedma, capital of Rio Negro, that was replaced for the urban-rural conglomerate of Alto Valle del Rio Negro. 6 Heckman et al (1998) find that the mismatch due to different questionnaire and different labor markets amounts to half of the selection bias in their analysis. 7 A follow-up survey of the same Trabajar participants was done in October/November 1999, to coincide with the next round of the PHS, and similarly in May/June 2000.7 The PHS has a rotating panel design with one quarter replaced each round, so it was possible to form a panel for the comparison group. Our matches were constrained to only include those who would be followed up. Naturally this limits the matching options - particularly so by the second follow- up survey, by which time only half of the original sample is re-surveyed. The PHS is a far shorter survey instrument than that used by Jalan and Ravallion (1999) for their single difference estimate of the impact on incomes of Trabajar participation. Since there are fewer observables in the data, the matching is unlikely to be as good. Results in the literature suggest that single-difference PSM estimates can be unreliable when the data available do not include important determinants of participation (Heckman et al., 1997, 1998; Smith and Todd, 2000). However, here we have the advantage that we can follow up participants over time, exploiting the rotating panel design of the PHS. Thus, although we cannot expect that our single difference PSM estimates will be as r eliable as in Jalan and Ravallion, we can eliminate the time-invariant errors due to miss-matching arising from violations of the conditional independence assumption. There was a sharp contraction is Trabajar participation after the first two surveys. 49% of the Trabajar workers interviewed in the baseline survey were no longer employed under the program in the first follow-up survey (Table 1). Only 16% of the original Trabajar workers were employed on the program by the second follow-up survey. This contraction in employment on the program did not appear to stem from a "pull" effect from the rest of the economy. There was little sign of economic recovery between the surveys. The overall unemployment rate increased 7 A fourth survey was done six months later. Over 90% of the initial Trabajar workers had left the program by the fourth wave. There were too few continuing participants to facilitate further analysis. 8 in one of the provinces (Chaco) and fell, but not greatly, in the other two; see Table 2, which also gives unemployment rates for the second follow up survey, six months later, and for six months prior to the first survey. The large number of participants leaving the program appears instead to be due to a normal process of rotation arising from the fact that projects do not last longer than six months. When a project ends, its beneficiaries are not incorporated automatically in another project. The responsible organizations are the ones that select the participants. In the country as a whole, 45% of Trabajar workers participate in only one project (46% in Chaco, 52% in Mendoza and 51% in Tucuman). On top of this designed rotation, there was a severe contraction in aggregate outlays on the program starting at the end of 1999. This was an outcome of overall fiscal austerity, to keep Argentina within macroeconomic targets. Aggregate spending on the program by the center in the first five months of 2000 was only 29% of its level in the last five months of 1999. Existing projects were completed, but the number of new projects approved shrank sharply in the latter part of 1999, to bring down the center's outlays. As already noted, the wage rate was also cut; Table 3 gives the sample mean wages by survey round. The aggregate cuts to the program made it less likely that past participants would find another project to join. A large new workfare program, the Emergency Employment program took up some of the slack in 2000. This was not in operation by the time of the second survey (first follow-up survey), but it was by the third survey. While our impact estimates using the first and second surveys are not likely to be affected by this new program, this is not true of the results using the third survey. 9 3. Estimation methods Our strategy is to compare income changes between those who stay in the program and those who leave, after netting out the income changes for an observationally similar comparison group of non-participants. This is an examnple of what has been called in the literature a "difference-in-difference-in-difference" or "triple-difference" estimate.8 We first discuss our method of selecting comparison groups, both of initial participants and for continuing participants. We then describe our version of the triple-difference estimator. 3.1 Controllingfor observed heterogeneity We use PSM to balance observed covariates at two stages. Firstly we form a matched comparison group for initial participants and secondly we match those who continue to participate over time ("stayers") with those that drop out ("leavers"). The second stage matching deals with the observed differences between subsequent leavers and stayers. PSM balances the distributions of covariates between participants and a comparison group based on similarity of their predicted probabilities of participation (their "propensity scores"). Rosenbaum and Rubin (1983) show that exact matching on the basis of propensity scores eliminates the bias in identifying the causal effect due to covariates. PSM is thus the observational analog of an experiment in which participation is independent of outcomes; the difference is that a pure experiment does not require the untestable assumption of independence conditional on observables. 8 The triple-difference method appears to have been first used by Gruber (1994) who included interaction effects between time and location (as well as separate time and location effects) in modeling the earnings effects of labor laws in the US. 10 Two groups are identified: those that participate (Di =1) and those that do not (D,=0). We rule out interference between units under the assumption that the gain to a worker from participation in a program such as Trabajar does not spillover to nonparticipants.9 Participants are matched to individuals who did not participate on the basis of the propensity score, defined as P(x,) = Pr(D, = 1lx1) where xi is a vector of pre-exposure control variables. Rosenbaum and Rubin (1983) prove that if the Di's are independent over all i, and outcomes are independent of participation given xi (i.e. unobserved differences do not influence whether or not i participates), then outcomes are also independent of participation given P(x1), just as they would be if participation was assigned randomly.'0 The value of P(x) is used to select control subjects for each of those treated. This eliminates bias in estimated treatment effects due to differences in the covariates. In practice the propensity score must be estimated. Here we follow the standard practice in PSM applications of using the predicted values from standard logit models to estimate the propensity score for each observation in the participant and the comparison-group samples. Using the estimated propensity score, matched-pairs are constructed on the basis of how close the scores are across the two samples. The nearest neighbor to the i'th participant is defined as the non-participant that minimizes [P(x,) - P(xj )]2 over allj in the set of non-participants, where P(Xk) is the predicted propensity score for observation k. Matches are only accepted if [P(xi) - P(x1)V is less than 0.00001 (an absolute difference less than 0.0032.) We only include those observations on non-participants that share a common range of values of the propensity 9 In the matching literature, this is the stable unit-treatment value assumption (Rosenbaum and Rubin, 1983). 10 The assumption that outcomes are independent of participation given xi is variously referred to in the literature as the "conditional independence," "strong ignorability," or "selection on observables." 11 scores calculated for the participants (i.e., the two groups share common support in the predicted propensity scores). In the following analysis, the comparison group for each participant is defined as the set of five nearest neighbors amongst non-participants in terms of the predicted propensity scores. 3.2 Latent heterogeneity PSM gives unbiased estimates of program impact if selection into the program is based solely on observables; selection bias due to (non-ignorable) latent heterogeneity will remain in single difference comparisons using PSM. With access to a pre-intervention baseline, one can eliminate time-invariant selection bias due to unobservables, by differencing over time. One can also deal with time-invariant selection bias with only post-intervention data, using follow-up surveys. However, in doing so one must recognize that the gains to current non-participants need not be symmetric before and after participation; while one may be happy to assume that baseline units are unaffected by the program, it is far less plausible that drop outs gain nothing currently from their past participation. So there are two distinct sources of selection bias in our set-up. One is in the existence of latent heterogeneity in who participates in the program, leading to miss- matching in determining the comparison group, and hence a systematic error in measuring the counter-factual income. Secondly, latent heterogeneity may affect the decision to stay in the program or drop out. Participants with high (unobserved) gains from participation may well be less likely to drop out of the program. This source of bias does not disappear. The double difference ("difference-in-difference" or DD) estimate is the difference in the income gains over time between a treatment group of program participants and a matched comparison group of non-participants. Our triple difference estimate (DDD) is defined as the difference between the value of the double difference for stayers and leavers. 12 Without loss of generality we can write the observed income of a Trabajar participant at date t as: yi,t=it+i (t21 1 where Yj, is the counter-factual income of the Trabajar participant if the program had not existed, and G1, is the income gain from participation (either participation at that date or previously). An indicator of the counter-factual income is available for a matched comparison group and is given by Yc. This is a noisy indicator due to miss-matching arising from latent heterogeneity. Thus the single difference estimator is potentially biased. We make the standard assumption in double-difference studies that the selection bias is time invariant, and so it is swept away by taking differences over time. More precisely, the first difference of Y~, is assumed to provide an unbiased estimate of the first difference of Yt,: E(AY,c) = AY,* (2) where A refers to the difference between the value at t and t- 1. From (1) and (2) it is evident that: E[A(Yj, - Yc,)] = AGi, (3) In the usual double difference set up, period 1 precedes the intervention, and it is assumed that Gil = 0 for all i. However, in our case, the program is in operation in period 1. The scope for identification arises from the fact that some participants at date 1 drop out of the program at date 2. Let Di, = 1 if individual i stays in the program, and let Di, = 0 if she does not. Our triple-difference estimator for t-2 is then: 13 DDD E- [A(Y,T - Yc )ID2 = 1] - E[A(YiT - yc )ID,2 =0= [E(GC2 IDi2 = 1) - E(Gi2 IDi2 = 0)] - [E(G1 ID,2 = 1) - E(Gj, IDi2 = 0)] (4) The first term in square brackets on the far RHS of (4) is the net gain to continued participation in the program, given by the difference between the gain to participants in period 2 and the gain to those who dropped out. This provides our measure of income replacement. Notice that there may be some gain from past participation for those who drop out (E(G,2 ID,2 = 0)
Группа Всемирного банка · Policy Research Working Paper
工作福利制的参与者能从紧缩中迅速恢复吗?
Открыть оригинал документа
Полный текст размещён на сайте публикующей организации. lawenc.com индексирует метаданные и ведёт на официальный источник.
Полный текст
Основные сведения
Организация
Группа Всемирного банка
Тип документа
Policy Research Working Paper
Страна
Аргентина
Источник
Всемирный банк