A statistical approach to address the problem of heaping in self-reported income data

A statistical approach to address the problem of heaping in self-reported income data
复制标题

解决自我报告收入数据堆积问题的统计方法

DOI:
10.1080/02664763.2015.1077372
复制
发表时间:
2015
影响因子:
1.5
通讯作者:
A. Würbach
A. Würbach
中科院分区:
数学4区
文献类型:
--
作者:
S. Zinn;A. Würbach

文献摘要

被引文献

相似文献

自我报告的收入信息特别受到故意粗化数据的影响,这被称为堆积或舍入。如果它不是完全随机发生的-通常情况下-堆积和四舍五入对统计分析的结果有不利影响。传统的统计方法没有考虑这种报告偏倚,因此可能会产生无效的推断。我们描述了一种新的统计建模方法,使我们能够以充分和灵活的方式处理自我报告的堆积收入数据。我们建议建模堆积机制和真正的底层模型相结合。为了描述真实的净收入分布,我们使用零膨胀对数正态分布。通过比较假设的收入分配和经验的启发式程序,从数据中确定堆积点。为了确定堆积行为,我们采用两种不同的模型:要么我们假设分段恒定的堆积概率,或堆积概率被认为是随着接近堆积点而稳步增加。我们验证了我们的方法通过一些例子。为了说明所提出的方法的能力,我们进行了一个案例研究,使用德国国家教育小组研究的收入数据。
Self-reported income information particularly suffers from an intentional coarsening of the data, which is called heaping or rounding. If it does not occur completely at random – which is usually the case – heaping and rounding have detrimental effects on the results of statistical analysis. Conventional statistical methods do not consider this kind of reporting bias, and thus might produce invalid inference. We describe a novel statistical modeling approach that allows us to deal with self-reported heaped income data in an adequate and flexible way. We suggest modeling heaping mechanisms and the true underlying model in combination. To describe the true net income distribution, we use the zero-inflated log-normal distribution. Heaping points are identified from the data by applying a heuristic procedure comparing a hypothetical income distribution and the empirical one. To determine heaping behavior, we employ two distinct models: either we assume piecewise constant heaping probabilities, or heaping probabilities are considered to increase steadily with proximity to a heaping point. We validate our approach by some examples. To illustrate the capacity of the proposed method, we conduct a case study using income data from the German National Educational Panel Study.