Efficient estimation in a partially specified nonignorable propensity score model

Efficient estimation in a partially specified nonignorable propensity score model
复制标题

DOI:
10.1016/j.csda.2021.107322
复制
发表时间:
2022-06-09
影响因子:
1.8
通讯作者:
Zhao,Jiwei
Zhao,Jiwei
中科院分区:
数学3区
文献类型:
--
作者:
Li,Mengyan;Ma,Yanyuan;Zhao,Jiwei

文献摘要

被引文献

相似文献

考虑回归设置,其中响应变量受到缺失数据的影响,而协变量则得到充分观察。一个不可否认的倾向评分模型,即,在所有变量条件下观察到响应的概率取决于缺失值本身,这在整个论文中都是假设的。在这类问题中,模型误设定和模型可辨识性是两个关键问题。一个完全参数的方法可以产生的结果是敏感的模型假设,而一个完全非参数的方法可能是不够的模型识别。提出了一种新的灵活的半参数倾向评分模型,其中缺失指标和部分观察到的反应之间的关系是完全不确定的,估计非参数,而缺失指标和完全观察到的协变量之间的关系是参数建模。建议的估计是通过半参数处理构造的,并被证明是半参数有效的。进行全面的模拟研究,以检查有限样本的估计性能。虽然朴素参数方法导致严重偏差的估计和差的覆盖结果,所提出的方法产生的估计与可忽略的有限样本偏差,也正确的推断结果。所提出的方法通过用于血液样品中的白蛋白水平的电子健康记录(EHR)数据应用进一步说明。实证分析表明,所提出的半参数倾向评分模型比纯参数模型更合理。所提出的方法可能是非常有用的,以揭示未知的和可能的非线性依赖的倾向评分模型的白蛋白水平,并建议实际使用。
Consider the regression setting where the response variable is subject to missing data and the covariates are fully observed. A nonignorable propensity score model, i.e., the probability that the response is observed conditional on all variables depends on the missing values themselves, is assumed throughout the paper. In such problems, model misspecification and model identifiability are two critical issues. A fully parametric approach can produce results that are sensitive to the model assumptions, while a fully nonparametric approach may not be sufficient for model identification. A new flexible semiparametric propensity score model is proposed where the relationship between the missingness indicator and the partially observed response is totally unspecified and estimated nonparametrically, while the relationship between the missingness indicator and the fully observed covariates is modeled parametrically. The proposed estimator is constructed via a semiparametric treatment and is proved to be semiparametrically efficient. Comprehensive simulation studies are conducted to examine the finite-sample performance of the estimators. While the naive parametric method leads to heavily biased estimator and poor coverage results, the proposed method produces estimator with negligible finite-sample biases and also correct inference results. The proposed method is further illustrated via an electronic health records (EHR) data application for the albumin level in the blood sample. The empirical analyses demonstrated that the proposed semiparametric propensity score model is more sensible than a purely parametric model. The proposed method could be very useful to uncover the unknown and possibly nonlinear dependence of the propensity score model to the albumin level, and is recommended for practical use.