SEMI-SUPERVISED INFERENCE: GENERAL THEORY AND ESTIMATION OF MEANS

SEMI-SUPERVISED INFERENCE: GENERAL THEORY AND ESTIMATION OF MEANS
复制标题

DOI:
10.1214/18-aos1756
复制
发表时间:
2019-10-01
影响因子:
4.5
通讯作者:
Cai, T. Tony
Cai, T. Tony
中科院分区:
数学1区
文献类型:
--
作者:
Zhang, Anru;Brown, Lawrence D.;Cai, T. Tony

文献摘要

被引文献

相似文献

我们提出了一个一般的半监督推理框架,主要关注总体均值的估计。在半监督设置中,通常存在一个未标记的协变量向量样本和一个由协变量向量和实值响应(“标签”)组成的标记样本。否则,该公式是“少假设”的,因为没有对数据的统计或函数形式施加主要条件。我们考虑了理想的半监督设置,其中有无限多的未标记样本可用,以及普通的半监督设置,其中只有有限数量的未标记样本可用。给出了总体均值的估计量和相应的置信区间。理论分析了该方法的渐近分布和l(2)-风险。令人惊讶的是,基于最小二乘法的简单形式提出的估计器优于普通样本均值。估计量的简单、透明的形式使我们相信,它对普通样本均值的渐近改进即使对于中等大小的样本也几乎成立。将该方法进一步推广到非参数条件下,该条件下的预测率可以渐近地得到。通过模拟研究和涉及无家可归人口估计的真实数据示例进一步说明了所提出的估计。
We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of covariate vectors along with real-valued responses ("labels"). Otherwise, the formulation is "assumption-lean" in that no major conditions are imposed on the statistical or functional form of the data. We consider both the ideal semi-supervised setting where infinitely many unlabeled samples are available, as well as the ordinary semi-supervised setting in which only a finite number of unlabeled samples is available.Estimators are proposed along with corresponding confidence intervals for the population mean. Theoretical analysis on both the asymptotic distribution and l(2)-risk for the proposed procedures are given. Surprisingly, the proposed estimators, based on a simple form of the least squares method, outperform the ordinary sample mean. The simple, transparent form of the estimator lends confidence to the perception that its asymptotic improvement over the ordinary sample mean also nearly holds even for moderate size samples. The method is further extended to a nonparametric setting, in which the oracle rate can be achieved asymptotically. The proposed estimators are further illustrated by simulation studies and a real data example involving estimation of the homeless population.