ESTIMATING INCOME STATISTICS FROM GROUPED DATA: MEAN-CONSTRAINED INTEGRATION OVER BRACKETS

ESTIMATING INCOME STATISTICS FROM GROUPED DATA: MEAN-CONSTRAINED INTEGRATION OVER BRACKETS
复制标题

DOI:
10.1177/0081175018782579
复制
发表时间:
2018-01-01
期刊:
SOCIOLOGICAL METHODOLOGY, VOL 48
影响因子:
--
通讯作者:
Wheeler, Christopher A.
Wheeler, Christopher A.
中科院分区:
其他
文献类型:
--
作者:
Jargowsky, Paul A.;Wheeler, Christopher A.

文献摘要

被引文献

相似文献

研究收入不平等、经济隔离和其他课题的研究人员必须经常依靠分组数据——也就是说,在这些数据中,成千上万的观察结果被简化为按特定收入等级计算的单位数。在这个等级内的家庭分布是未知的,最高收入通常包括在一个开放式的最高等级,比如“20万美元及以上”。该估计问题的常用方法包括计算中点估计,并在上括号中假设Pareto分布,并拟合数据的灵活多参数分布。作者描述了一种新的方法,即平均约束的括号积分(MCIB),它比那些只使用括号计数和数据总体平均值的方法要准确得多。在对297个大都市区进行分析的基础上,MCIB得出了标准差、基尼系数和泰尔指数的估计值,这些估计值分别为0.997、0.998和0.991,与从基础个人记录数据中计算出来的参数相关。分配的百分位数和分配的五分位数所占的收入份额也获得了类似水平的准确性。该方法可以很容易地推广到其他分布参数和不等式统计中。
Researchers studying income inequality, economic segregation, and other subjects must often rely on grouped data-that is, data in which thousands or millions of observations have been reduced to counts of units by specified income brackets. The distribution of households within the brackets is unknown, and highest incomes are often included in an open-ended top bracket, such as "$200,000 and above." Common approaches to this estimation problem include calculating midpoint estimators with an assumed Pareto distribution in the top bracket and fitting a flexible multiple-parameter distribution to the data. The authors describe a new method, mean-constrained integration over brackets (MCIB), that is far more accurate than those methods using only the bracket counts and the overall mean of the data. On the basis of an analysis of 297 metropolitan areas, MCIB produces estimates of the standard deviation, Gini coefficient, and Theil index that are correlated at 0.997, 0.998, and 0.991, respectively, with the parameters calculated from the underlying individual record data. Similar levels of accuracy are obtained for percentiles of the distribution and the shares of income by quintiles of the distribution. The technique can easily be extended to other distributional parameters and inequality statistics.