A statistical approach to address the problem of heaping in self-reported income data
A statistical approach to address the problem of heaping in self-reported income data
复制标题
解决自我报告收入数据堆积问题的统计方法
DOI:
10.1080/02664763.2015.1077372
复制
发表时间:
2015
影响因子:
1.5
通讯作者:
A. Würbach
中科院分区:
文献类型:
--
作者:
S. Zinn;A. Würbach
Self-reported income information particularly suffers from an intentional coarsening of the data, which is called heaping or rounding. If it does not occur completely at random – which is usually the case – heaping and rounding have detrimental effects on the results of statistical analysis. Conventional statistical methods do not consider this kind of reporting bias, and thus might produce invalid inference. We describe a novel statistical modeling approach that allows us to deal with self-reported heaped income data in an adequate and flexible way. We suggest modeling heaping mechanisms and the true underlying model in combination. To describe the true net income distribution, we use the zero-inflated log-normal distribution. Heaping points are identified from the data by applying a heuristic procedure comparing a hypothetical income distribution and the empirical one. To determine heaping behavior, we employ two distinct models: either we assume piecewise constant heaping probabilities, or heaping probabilities are considered to increase steadily with proximity to a heaping point. We validate our approach by some examples. To illustrate the capacity of the proposed method, we conduct a case study using income data from the German National Educational Panel Study.