EFFICIENT AND ADAPTIVE LINEAR REGRESSION IN SEMI-SUPERVISED SETTINGS

EFFICIENT AND ADAPTIVE LINEAR REGRESSION IN SEMI-SUPERVISED SETTINGS
复制标题

DOI:
10.1214/17-aos1594
复制
发表时间:
2018-08-01
影响因子:
4.5
通讯作者:
Cai, Tianxi
Cai, Tianxi
中科院分区:
数学1区
文献类型:
--
作者:
Chakrabortty, Abhishek;Cai, Tianxi

文献摘要

被引文献

相似文献

我们考虑半监督环境下的线性回归问题,其中可用数据通常包括:(I)小的或中等大小的有标签的数据,和(Ii)大得多的大得多的无标签数据。与协变量不同,这些数据自然产生于结果获取成本高昂的环境中,这是现代研究中涉及电子医疗记录(EMR)等大型数据库的常见场景。像普通最小二乘(OLS)估计器一样的监督估计器只利用标记数据。在所采用的线性模型中,是否以及何时可以利用未标记数据来改进回归参数的估计一直是人们感兴趣的问题。本文提出了一类“高效和自适应半监督估计”(EASE)来提高估计效率。容易的是适应模型错误指定的两步估计器,在模型错误指定的情况下导致改进(在某些情况下是最优的)效率,在线性模型下导致相同(最优的)效率。这种自适应特性在现有文献中往往没有被提及,它对于倡导“安全”使用未标记的数据至关重要。EASY的构建主要包括灵活的“半非参数”归因,包括一个即使在协变量数量不少的情况下也能很好地工作的平滑步骤;以及后续的“重新调整”步骤和交叉验证(CV)策略,这两个步骤都对解决两个重要问题具有有用的实践和理论意义:欠平滑和过拟合。我们建立了渐近结果,包括相合性、渐近正态和EASE的自适应性质。我们还提供了影响函数展开和用于推理的“双”CV策略。通过广泛的模拟进一步验证了结果,随后将其应用于关于自身免疫的EMR研究。
We consider the linear regression problem under semi-supervised settings wherein the available data typically consists of: (i) a small or moderate sized "labeled" data, and (ii) a much larger sized "unlabeled" data. Such data arises naturally from settings where the outcome, unlike the covariates, is expensive to obtain, a frequent scenario in modern studies involving large databases like electronic medical records (EMR). Supervised estimators like the ordinary least squares (OLS) estimator utilize only the labeled data. It is often of interest to investigate if and when the unlabeled data can be exploited to improve estimation of the regression parameter in the adopted linear model.In this paper, we propose a class of "Efficient and Adaptive Semi-Supervised Estimators" (EASE) to improve estimation efficiency. The EASE are two-step estimators adaptive to model mis-specification, leading to improved (optimal in some cases) efficiency under model mis-specification, and equal (optimal) efficiency under a linear model. This adaptive property, often unaddressed in the existing literature, is crucial for advocating "safe" use of the unlabeled data. The construction of EASE primarily involves a flexible "semi-nonparametric" imputation, including a smoothing step that works well even when the number of covariates is not small; and a follow up "refitting" step along with a cross-validation (CV) strategy both of which have useful practical as well as theoretical implications towards addressing two important issues: under-smoothing and over-fitting. We establish asymptotic results including consistency, asymptotic normality and the adaptive properties of EASE. We also provide influence function expansions and a "double" CV strategy for inference. The results are further validated through extensive simulations, followed by application to an EMR study on auto-immunity.