WPS I9 v2 POLICY RESEARCH WORKING PAPER 1902 When Economic Reform Household surveydaii income Inequality incre,Clasn is Faster than Statistical in post-reform rurlra Reform but this may re lect used to process car[a Fathe: than the real efre(t of Measuring and Explaining Inequality structural chancies on * T 1 r>1 * ~~~~~~~~~~~~~rural economy. in Rural China 1\Martiz kRz'alliomz Shblobua Cheni The World Bank Development Research Group March 1998 I POLICY RESEARCH WORKING PAPER 1902 Summary findings Official tabulations from household survey data suggest allowances are made for regional cost-of-living rising income inequality in post-reform rural China, a differences. trend of public concern. But the structural changes in The data revisions also suggest somewhat different China's rural economy have not been properly reflected explanations for rising inequality. Nonfarm income was in the methods used to process raw survey data. secondary to grain production. While access to farm land Using micro data for four provinces, Ravallion and was relatively equal, higher returns to land over time Chen find that two-thirds of the conventionally were inequality-increasing. But holding other factors measured increase in inequality in 1985-90 vanishes constant, lower returns to physical capital reduced when market-based valuation methods are used and inequality over time, as did private transfers. This paper - a product of the Development Research Group - is part of a larger effort in the group to improve data on poverty and inequality in developing countries. The study was funded by the Bank's Research Support Budget under the research project "Dynamics of Poverty in Rural China." Copies of this paper are available free from the World Bank, 1818 H Street NW, Washington, DC 20433. Please contact Patricia Sader, room MC3-632, telephone 202-473-3902, fax 202- 522-1153, Internet address psader@worldbank.org. March 1998. (38 pages) The Policy Researcb Working Paper Sedbes disseminates the fiydReags of work in progress to encourage the exchange of ideas aborCt development issufes. An objective of the series is to get the findings otit quickly, even if the presentations are less than fully polished. The papers carry the names of the authors and should be cited accordinigly. The findings, interpretations, and conclusions expressed in this paper are entirely those of the authors. They do not necessarily represent the view of the World Bank, its Execaftive Directors, or the |countries they represent. Prod-uced by the Policy Research Dissemination Center When Economic Reform is Faster than Statistical Reform: Measuring and Explaining Income Inequality in Rural China Martin Ravallion and Shaohua Chen 1 Introduction There is a widespread view that the transition from a socialist economic system to a market economy will entail rising inequality, and there is support for that view in recent compilations of distributional data for the 1980s and '90s (Milanovic, 1996; Ravallion and Chen, 1997). However, these compilations are typically based on the tabulations of distributional data (drawn from household surveys) that are made available by official sources. While economic reforms often have important implications for the methods used in measuring economic welfare and inequality, government statistical agencies may not be adjusting as rapidly as one would like to the structural changes going on in the economy. And users of the official data rarely probe into the raw micro data underlying the distributional comparisons being made, either because of lack of access to the data or lack of resources for doing so. Could lags in reforming statistical methods entail substantial biases in assessments of how inequality is changing during the transition? The structural changes going on are not necessarily inequality-increasing. A common element of socialist economic planning was the suppression of food-staple prices, to help finance industrialization.2 Through market liberalization, the transition typically entails higher food staple prices. To the extent that food-staple producers are concentrated among the poor, the transition will put downward pressure on inequality. If all incomes were derived from market exchange then these effects should be seen quickly in official data on distribution drawing on household surveys. However, a large share of income in poor rural economies takes the form of direct consumption of own production. Valuations must be 2 This vvas often referred to as the "price scissors" and there is a large literature on the practice; for a recent analysis and references see Sah and Stiglitz (1992). 2 imputed for this and other income sources which were not acquired through exchange. When prices are controlled by administrative fiat, the same prices are naturally used for valuation. But there can be no assurance that old administrative prices will be replaced by market prices as the transition proceeds. Unless statistical agencies are quick enough to adapt to such changes, biases can enter survey-based analyses of (among other things) income inequality. The transition can have many other implications for measurement. The level of prices may rise faster in some regions of the economy than others after reforms (reflecting nontraded goods, or less than perfect spatial market integration, due for instance to poorly developed transportation). If it were the initially better-off regions which saw higher growth and higher inflation (due to higher aggregate demand locally) then assessments of income distribution which ignored geographic differences in prices could overestimate the rate at which inequality was increasing. There is no good a priori reason to assume that there will be a bias, or that (when there is) it could go only one way. For example, the share of income from undervalued components may be no different between the rich and poor, or the rate of inflation may be higher in poorer regions. These are empirical questions, although they can be difficult to answer since they require access to, and reprocessing of, the raw data underlying official tabulations. This paper addresses these concerns in the context of post-reform rural China. Beginning with Premier Deng's reforms in 1978, China's rural economy became market-oriented; prices were freed and the farm-household replaced the commune as the decision-making unit. These reforms brought about changes to data collection, including greater reliance on household surveys. The scope and collection methods of such surveys improved significantly during the 1980s, starting with the Rural Household Survey (RHS) introduced in 1984. This has been the main source of 3 data for distributional analysis on rural China. Tabulations of results from the RHS in China's Statistical Yearbooks have suggested rising income inequality since the mid-i 980s. This has been widely reported and attracted much attention? However, there are reasons to be cautious in interpreting the available evidence on income inequality in rural China. A number of potential problems have been identified in recent literature, including the undervaluation of income in kind from the consumption of own-farm products due to continuing reliance on planning prices for valuation purposes.4 We examine how the problems in official tabulations from the household survey data have affected measurements of the overall level of inequality, and how it changes over time. We also examine how these data problems impinge on explanations of the observed changes in overall inequality.5 Suppose, for example, that we want to know if the rising income inequality in China is due to the booming rural non-farm sector (including the famous Township and Village Enterprises). Or we may want to see what role public and private transfers played. In principle, the answers to such questions will depend on the method used to measure incomes at the household level. For example, undervaluing income in kind from own production might lead one to underestimate the contribution of this income component to rising income inequality, given that 3 See, for example, the front page article in The New York Times, December 27, 1995. ' Discussions of the problems can be found in World Bank (1992), Khan et al. (1993) and Chen and Ravallion (1996). 5 There have been a number of studies attempting to throw light on the causes of inequality in China since reforms began in the late 1970s. Decompositions have been done along various dimensions (geographic and by income source) and at various levels of spatial aggregation (some by county, some by village, some household) and for differing time dimensions (some using single cross-sectional surveys, some including comparisons over time). Contributions include Knight and Song (1993), Rozelle (1994) and Howes and Hussain (1994). 4 its progressive undervaluation over time would probably lead one to conclude (incorrectly) that this income component is becoming less covariate with total income. It is an empirical question just how robust explanations of rising inequality are to these data problems. We address these issues using a large household-level data set for rural China spanning the period 1985-90. The region we study embraces booming Guangdong on the coast (the province surrounding Hong Kong) and the far less prosperous, and more economicly stagnant, inland provinces of Guangxi, Yunnan and Guizhou. Having access to the micro data means that we can attempt to correct the main concerns about existing distributional data. After making corrections to the processing, we are able to use the survey to address a number of questions about the proximate causes of the observed changes in income inequality. The following section summarizes the theoretical results we will be using from the literature on inequality measurement. Section 3 then looks at the theoretical implications of undervaluing an income component for measures of inequality and their decomposition. In section 4 we describe our data, while section 5 gives our overall results on income inequality, with and without our corrections to the data processing. We then turn in section 6 to the task of explaining the observed changes in inequality. Our conclusions are summarized in section 7. 2 Inequality measurement and decomposition methods A measure of inequality can be written in generic form: I = I(y 1/A Y2/R . ...YNIR )(1 where y, is the i'th person's income in a population of size N, and 1 is mean income. We assume 5 that this measure is continuous, symmetric (swapping incomes does not change the measure), normalized such that inequality is zero when all persons have the same income, and that the measure satisfies the "transfer axiom" such that a transfer from rich to poor reduces inequality. For some sorts of distributional comparisons we may not need to know any more about the measure of inequality. For example, if the Lorenz curve (giving, on the vertical axis, the share of total income held by the poorest x% of the population) for distribution A is everywhere above that of B then all inequality measures in the above class of measures will show higher inequality in B than A (Atkinson, 1970). In our empirical work we will focus on two special cases of the above class of measures. The first is the well-known Gini index (G), given by the (household-size weighted) mean absolute deviation between all pairs of per capita household incomes. The second is a member of the Generalized Entropy class of additively decomposable measures, namely the average log deviation of incomes from their mean:6 LD = - log(i /Y1) (2) N We will also be interested in explaining inequality and its changes over time. There are potentially many ways of decomposing a change in inequality by income source. Here we follow a strand of the theoretical literature which has constrained the choices by postulating certain 6 If N stood instead for the number of households then household-size weights would appear in this formula. All statistics in this paper which are based on the household-level data have been household-size weighted. 6 axioms that are deemed desirable for any decomposition. (We only summarize the basic results that will be needed for the empirical work later.) Let total income (per person) be divided into m categories, such that, for the i'th household: m Yi kE Yik (3) k= 1 If these components were uncorrelated with each other, and one measured inequality by the squared coefficient of variation (CV), then the natural decomposition would be to measure the contribution of each income component to inequality by its squared CV. However, in practice different income sources are correlated to varying extents. And there are many other inequality measures that one might want to consider besides squared CV. How then should one apportion total inequality between components? A powerful result proved by Shorrocks (1982) shows that a modified version of the squared-CV decomposition (allowing for non-zero correlations) can also be defended as a decomposition method for a wide range of inequality measures. For the class of inequality measures described above,7 Shorrocks shows that the proportion of total inequality contributed by the k'th income source is given by: cov(yV,y) rks Ck = 4=k k var(y) s ' In fact Shorrocks proves the following result for an even larger class of measures; see his paper for full details. 7 where rk is the correlation coefficient with total income and Sk and s are the standard deviation of the k'th income component and of total income respectively. Note that Ck sums to one over all k and is simply the ordinary least squares regression coefficient of Yk ony. The decomposition based on (4) is independent of the precise measure of inequality used (within the aforementioned class of measures). Notice that the contribution of any income component to total inequality depends on both the variance of that component (relative to the variance in total income) and its correlation coefficient with total income. So the fact that some income component contributes a lot to total inequality does not necessarily mean that it is itself very unequally distributed; it may instead be highly correlated with total income, yet quite equally distributed. Similarly, a highly unequally distributed income component may contribute little to total inequality because it is roughly uncorrelated with total income, or it may be inequality-reducing because of a negative correlation with total income. The above result holds for a decomposition of the level of inequality. What about changes in inequality over time? Building on the Shorrocks' decomposition, we follow Jenkins (1995) and Fields (1996) in calculating the contribution of the k'th income source to the change in total inequality between dates 1 and 2 by: k2I2 kl I(5) 12 -II I2 1, which sums to one. Notice that (unlike the levels decomposition) this decomposition will depend on the specific inequality measure used. We will compare results for the Gini index with those for 8 the average log deviation given by (2). One can also ask how much of the level of inequality or its change over time is due to some variable determining income through a stochastic process. To do so, replace equation (3) with a regression model for income: m Y= 3kik (6) where xik is the k'th asset (xi. can be taken to be an error term, with pm=l). Following Fields (1996), the contribution of the k'th explanatory variable to total inequality is given by: Pkcov(Xk, Y) k var(y) This is simply the product of the partial regression coefficient of income on schooling (holding all other variables constant) with that total regression coefficient of schooling on income (holding nothing else constant). The contributions of each asset to the changes over time can then be determined using equation (5). The precise decomposition will naturally depend on the regression specification in (6). This should be borne in mind when interpreting the results. 3 Effects of valuation errors on measured inequality and its decomposition It is known that inequality measures can be highly sensitive to measurement errors; a few bad observations can have a large impact on measured inequality (Cowbell and Victoria-Fester, 1996). Here we are concerned with a particular structure of measurement error, arising from 9 undervaluation of an income component, as discussed in the introduction. We cannot find a treatment of this case in the literature, so we offer some observations, to help interpret the empirical results later. We examine effects of undervaluation on the level of inequality, the factor decomposition of inequality, and on the decomposition of changes in inequality over time. Let us first consider the effect of the valuation error on the level of measured inequality, as this is the easiest case. The revaluation can be thought of as a negative income tax. Let the average rate of revaluation (analogous to the average tax rate) be defined as the increase in imputed value as a proportion of original income. Following results from the literature on tax progressivity (see, for example, Pfingsten, 1988), the correction for undervaluation will lead to lower (higher) measured inequality if the average rate of revaluation falls (rises) as income increases. What about the effect on the factor decomposition of inequality? Recall that the share of inequality attributed to a given income component is the regression coefficient of that component on total income (equation 4). Both the regressor and regressand are underestimated (by the same amount). There will be two sources of bias in the regression coefficient; the first is the usual attenuation bias due to miss-measurement of the regressor, while the second is the bias due to the fact that the same error contaminates the regressand. These two biases will work in opposite directions and so one cannot say on a priori grounds what effect this will have on the regression coefficient. Intuitively, the lower the regression coefficient, the less important will be the second source of bias. So one expects undervaluation to lead to underestimation of the contribution to inequality when that contribution is sufficiently low. We can derive a very simple sufficient condition for signing the effect when the k'th 10 income component is undervalued by a constant proportion, such that the revaluation yields: Yk = (1 + a)yk (8) for a>0. We assume that 1 > ck> 0, although this can be relaxed; the following result holds for 1 +1/a> c > -(1 + a2vd/(2a)wherevk -var(ydlvar(y). On revaluing the undervalued component, its contribution to total inequality becomes: COV(Yk'Y) ( +a)(Ck + avd kV var(y ) +a vk + 2ac From (9) it is readily verified that c** > ck if and only if k k (2ck - I)ck k I + a(l -ck) So a sufficient condition for the undervaluation to underestimate the contribution to inequality is that the undervalued component of income accounts for less than one half of inequality. The effect of undervaluation on a factor's contribution to changes in inequality over time (yk given by equation 5) is more complicated, since it will clearly also depend on how the factor decomposition evolves. We confine attention to the case of empirical interest later in which inequality is increasing (with or without revaluation) and the undervalued income component's 11 contribution to inequality is underestimated. Let * denote the contribution of the k'th income component to rising inequality. It is readily verified that: * (kl Ck;) + (C -k2)I2],I* + [(Ck; - ckYl) + (ck2 - k;)2]22 Tk 7yk (I2 - I1)(I2* -I,) (11) If the factor decomposition does not change over time (c c * and c =c ) then clearly k2 ki k2 Ckl)te lal * = ct -- C*] = c - c* < ;the undervaluation ofthe k'th income component also leads to Yk-'k = ki .Cki =k2 k2 an underestimation of its contribution to rising inequality. However, the outcome is ambiguous when the factor decomposition is changing over time. From (11), the sign of - will also depend on the "cross-terms", c ck2 and ck2 c at least one of which must be positive.! A sufficient condition for ye > yk is that: c -c I c k2 kI < 2 Ckl kI (12) c * - cII * -c| k2 k2 k k2l However, it is entirely possible for revaluation to diminish the contribution of the undervalued income component to rising inequality, even when revaluation yields higher inequality at any one s The cross terms cannot both be negative, for then (c ckl) + (c2 - ck2) < 0 - a contradiction. 12 date. Suppose, for example, that with its undervaluation the measured contribution to inequality of the k'th income component does not change over time (ck2 =Cki ), but with the revaluation its contribution is found to fall over time (Ck* < ck i). Then * > y if and only if I2*/I* > (Ckl -CkY)(Ck; ck2) 4 Data The data are the household-level data from the Rural Household Survey (RHS) done by China's State Statistics Bureau (SSB). Our sample covers 9,500 rural households in Guangxi, Yunnan, Guizhou and Guangdong. The survey and steps we have taken in data processing are described in detail in Chen and Ravallion (1996). The RHS is a high quality survey in many respects, including both sampling methods and the care taken to minimize nonsampling errors through close supervision and regular visits to the sampled households. There are, however, problems in the methods used in processing the data after its collection, leading up to the tabulations found in the Statistical Yearbook for China. Chen and Ravallion (1996) review the main concerns about these data. We attempt to resolve the main problems by reprocessing the primary data for 1985-90. An important concern about the official data is that they continued to rely on old planning prices for the valuation of income-in-kind from consumption of own-farm production. These prices were below market prices (and also below government procurement prices). This undervalued a large component of income - notably non-marketed home production of grain - and 13 at a rising rate over time (Chen and Ravallion, 1996). The standard definitions indicate that, for our data set, an average of 21% of income came from grain production, of which 80% was the imputed value of consumption from own production. Other components of farm income appear also to have been undervalued, but this is a less worrying since the shares of income involved are smaller (22% of income came from non-grain farm output, but only 10% of this was from own consumption). Another problem is that the incomes used in past work have not included imputed rents for housing and durables. Past work has also ignored spatial differences in the cost of living. To deal with these problems, we have revalued grain-income in kind at median local (county-level) selling prices for grain, as determined from the primary household-level data.9 The administrative prices conventionally used by SSB for valuation were 72% of median selling price in 1985, and this had fallen to 48% by 1990. We have also imputed rents for housing and consumer durables, based on the asset valuations available in the primary survey data; we used five percent of the recorded dwelling value for housing and 10 percent for durables (Chen and Ravallion, 1996). And we have constructed new province-level spatial and inter-temporal cost of living indices. The spatial cost of living adjustment is based on poverty lines aiming to measure the local cost in 1988 of the same standard of living everywhere, based on a common bundle of foods and an allowance for non-food spending consistent with spending behavior at the food poverty line. The inter-temporal price indices are based on the rural CPI, though we have changed the weights to accord with consumption behavior of the poorest 30% of the population. Full details on both the poverty lines and the intertemporal cost-of-living deflators can be found in Chen and 9 Similar data are unavailable for revaluing other components of income in kind from own-farm production, although (as noted above) these appear to be minor. 14 Ravallion (1996).'
Группа Всемирного банка · Policy Research Working Paper
当经济改革快于统计改革时:测量和解释中国农村的不平等
Открыть оригинал документа
Полный текст размещён на сайте публикующей организации. lawenc.com индексирует метаданные и ведёт на официальный источник.
Полный текст
Основные сведения
Организация
Группа Всемирного банка
Тип документа
Policy Research Working Paper
Страна
Китай
Источник
Всемирный банк