Combining panel data sets with attrition and refreshment samples

Combining panel data sets with attrition and refreshment samples
复制标题

DOI:
10.1111/1468-0262.00260
复制
发表时间:
2001-11-01
期刊:
影响因子:
6.1
通讯作者:
Rubin, DB
Rubin, DB
中科院分区:
经济学1区
文献类型:
--
作者:
Hirano, K;Imbens, GW;Rubin, DB

文献摘要

被引文献

相似文献

在许多领域,研究人员希望考虑统计模型,这些模型允许比仅使用横截面数据推断出的更复杂的关系。在不同时间点重复观察相同单位的面板数据或纵向数据通常可以提供此类模型所需的更丰富的数据。 尽管此类数据使研究人员能够识别比横截面数据更复杂的模型,但面板中的数据缺失问题可能更为严重。特别是,即使在面板的初始波中做出响应的单位也可能在后续波中退出,使得具有面板的所有波的完整数据的子样本可能比原始样本更不具有总体代表性。有时,为了减轻磨损的影响而不丧失面板数据相对于横截面的优势,面板数据集通过用从原始总体中随机抽样的新单位替换已退出的单位来增强。 Ridder (1992) 使用这些替换单元来测试某些模型的磨损情况,我们将此类附加样本称为刷新样本。 我们探讨了这些样本对于估计损耗模型的好处。我们描述了刷新样本的存在允许研究人员测试面板数据中的各种损耗模型的方式,包括基于丢失数据随机丢失的假设的模型(MAR,Rubin,1976;Little 和 Rubin,1987)。 论文的主要结果精确地说明了刷新样品对消耗过程的信息程度;如果有更新样本,则可以识别一类不可忽略的缺失数据模型,而无需做出强有力的分布或函数形式假设。
In many fields researchers wish to consider statistical models that allow for more complex relationships than can be inferred using only cross-sectional data. Panel or longitudinal data where the same units are observed repeatedly at different points in time can often provide the richer data needed for such models. Although such data allows researchers to identify more complex models than cross-sectional data, missing data problems can be more severe in panels. In particular, even units who respond in initial waves of the panel may drop out in subsequent waves, so that the subsample with complete data for all waves of the panel can be less representative of the population than the original sample. Sometimes, in the hope of mitigating the effects of attrition without losing the advantages of panel data over cross-sections, panel data sets are augmented by replacing units who have dropped out with new units randomly sampled from the original population. Following Ridder (1992), who used these replacement units to test some models for attrition, we call such additional samples refreshment samples. We explore the benefits of these samples for estimating models of attrition. We describe the manner in which the presence of refreshment samples allows the researcher to test various models for attrition in panel data, including models based on the assumption that missing data are missing at random (MAR, Rubin, 1976; Little and Rubin, 1987). The main result in the paper makes precise the extent to which refreshment samples are informative about the attrition process; a class of non-ignorable missing data models can be identified without making strong distributional or functional form assumptions if refreshment samples are available.