On semi-supervised linear regression in covariate shift problems

On semi-supervised linear regression in covariate shift problems
复制标题

协变量平移问题中的半监督线性回归

DOI:
10.5555/2789272.2912101
复制
发表时间:
2015
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
M. Culp
M. Culp
中科院分区:
--
文献类型:
--
作者:
K. J. Ryan;M. Culp

文献摘要

被引文献

相似文献

半监督学习方法使用完整的训练(标记)数据和可用的测试(未标记)数据进行训练。使用未标记数据进行训练的价值的证明通常取决于将条件期望与边缘分布的高密度区域相关联的平滑性假设以及标记的完全随机假设的固有缺失。所谓的协变量转变对许多现有的半监督或监督学习技术提出了挑战。协变量平移模型允许标记和未标记特征数据的边缘分布不同,但给定特征数据的响应的条件分布是相同的。当依次获得完整的标记数据样本和未标记样本时,就会出现这种情况,因为很可能样本之间的特征数据分布差异很大。在这种实际的协变量移位问题中,在弹性网络训练期间使用未标记数据的价值在几何上是合理的。该方法的工作原理是获得未标记预测的调整系数,重新校准监督弹性网络以妥协:(i) 维持对标记数据的弹性网络预测,(ii) 将未标记预测缩小到零。当未标记特征数据集中在远离标记数据的低维流形上并且真实系数向量强调远离该流形的方向时,我们的方法被证明在未标记响应预测上占主导地位。当未标记响应预计较小时,未标记集上的监督预测的大方差减少的程度大于平方偏差的增加,因此偏差-方差权衡中的改进折衷是这种性能改进的理由。性能通过模拟和真实数据进行验证。
Semi-supervised learning approaches are trained using the full training (labeled) data and available testing (unlabeled) data. Demonstrations of the value of training with unlabeled data typically depend on a smoothness assumption relating the conditional expectation to high density regions of the marginal distribution and an inherent missing completely at random assumption for the labeling. So-called covariate shift poses a challenge for many existing semi-supervised or supervised learning techniques. Covariate shift models allow the marginal distributions of the labeled and unlabeled feature data to differ, but the conditional distribution of the response given the feature data is the same. An example of this occurs when a complete labeled data sample and then an unlabeled sample are obtained sequentially, as it would likely follow that the distributions of the feature data are quite different between samples. The value of using unlabeled data during training for the elastic net is justified geometrically in such practical covariate shift problems. The approach works by obtaining adjusted coefficients for unlabeled prediction which recalibrate the supervised elastic net to compromise: (i) maintaining elastic net predictions on the labeled data with (ii) shrinking unlabeled predictions to zero. Our approach is shown to dominate linear supervised alternatives on unlabeled response predictions when the unlabeled feature data are concentrated on a low dimensional manifold away from the labeled data and the true coefficient vector emphasizes directions away from this manifold. Large variance of the supervised predictions on the unlabeled set is reduced more than the increase in squared bias when the unlabeled responses are expected to be small, so an improved compromise within the bias-variance tradeoff is the rationale for this performance improvement. Performance is validated on simulated and real data.