Semisupervised inference for explained variance in high dimensional linear regression and its applications

Semisupervised inference for explained variance in high dimensional linear regression and its applications
复制标题

DOI:
10.1111/rssb.12357
复制
发表时间:
2020-01-20
影响因子:
5.8
通讯作者:
Guo, Zijian
Guo, Zijian
中科院分区:
数学1区
文献类型:
--
作者:
Cai, T. Tony;Guo, Zijian

文献摘要

被引文献

相似文献

本文考虑了半监督环境下高维线性模型 Y=X beta+epsilon 下解释方差 beta T sigma beta 的统计推断,其中 beta 是回归向量,sigma 是设计协方差矩阵。提出了一种校准估计器,它可以有效地集成标记和未标记数据。结果表明,估计器在一般半监督框架中实现了极小极大最优收敛速度。最优性结果表征了未标记数据如何对估计精度做出贡献。此外,建立了所提出的估计量的极限分布,并且未标记的数据也被证明有助于减少解释方差的置信区间的长度。所提出的方法被扩展到未加权二次函数 ||beta||22 的半监督推理。然后将获得的推理结果应用于一系列高维统计问题,包括信号检测和全局测试、预测精度评估和置信球构建。通过模拟研究和对具有多个性状的酵母分离数据集的遗传力估计分析,证明了合并未标记数据的数值改进。
The paper considers statistical inference for the explained variance beta T sigma beta under the high dimensional linear model Y=X beta+epsilon in the semisupervised setting, where beta is the regression vector and sigma is the design covariance matrix. A calibrated estimator, which efficiently integrates both labelled and unlabelled data, is proposed. It is shown that the estimator achieves the minimax optimal rate of convergence in the general semisupervised framework. The optimality result characterizes how the unlabelled data contribute to the estimation accuracy. Moreover, the limiting distribution for the proposed estimator is established and the unlabelled data have also proved useful in reducing the length of the confidence interval for the explained variance. The method proposed is extended to semisupervised inference for the unweighted quadratic functional ||beta||22. The inference results obtained are then applied to a range of high dimensional statistical problems, including signal detection and global testing, prediction accuracy evaluation and confidence ball construction. The numerical improvement of incorporating the unlabelled data is demonstrated through simulation studies and an analysis of estimating heritability for a yeast segregant data set with multiple traits.