Leveraging mixed and incomplete outcomes via reduced-rank modeling

Leveraging mixed and incomplete outcomes via reduced-rank modeling
复制标题

DOI:
10.1016/j.jmva.2018.04.011
复制
发表时间:
2018-09
期刊:
J. Multivar. Anal.
影响因子:
--
通讯作者:
Chongliang Luo;Jian Liang;Gen Li;Fei Wang;Changshui Zhang;D. Dey;Kun Chen
Chongliang Luo;Jian Liang;Gen Li;Fei Wang;Changshui Zhang;D. Dey;Kun Chen
中科院分区:
其他
文献类型:
--
作者:
Chongliang Luo;Jian Liang;Gen Li;Fei Wang;Changshui Zhang;D. Dey;Kun Chen

文献摘要

被引文献

相似文献

具有可能高维的多变量特征的多变量结果通常在各个领域产生。在许多现实问题中,收集的结果是混合类型的,包括连续测量,二元指标和计数,并且很大一部分值也可能缺失。无论其类型如何,这些混合的结果往往是相互关联的,代表了相同的基本数据生成机制的不同反映或观点。因此,综合多变量模型可能是有益的。我们开发了一个混合结果的降秩回归,它有效地实现了不同预测任务之间的信息共享。我们的方法集成了混合和部分观察到的结果属于指数分散家庭,通过假设所有的结果是通过一个共享的低维子空间跨越的功能。提出了一个一般的奇异值正则化准则,我们建立了一个非渐近性能界的监督学习的背景下,从指数族的混合结果和一般的抽样计划下的缺失数据的估计。提出了一种保证收敛的迭代奇异值阈值优化算法。我们的方法的有效性证明了模拟研究和应用预测健康相关的结果在纵向研究老化。
Multivariate outcomes with multivariate features of possibly high dimension are routinely produced in various fields. In many real-world problems, the collected outcomes are of mixed types, including continuous measurements, binary indicators and counts, and a substantial proportion of values may also be missing. Regardless of their types, these mixed outcomes are often interrelated, representing diverse reflections or views of the same underlying data generation mechanism. As such, an integrative multivariate model can be beneficial. We develop a mixed-outcome reduced-rank regression, which effectively enables information sharing among different prediction tasks. Our approach integrates mixed and partially observed outcomes belonging to the exponential dispersion family, by assuming that all the outcomes are associated through a shared low-dimensional subspace spanned by the features. A general singular value regularized criterion is proposed, and we establish a non-asymptotic performance bound for the proposed estimators in the context of supervised learning with mixed outcomes from an exponential family and under a general sampling scheme of missing data. An iterative singular value thresholding algorithm is developed for optimization with convergence guarantee. The effectiveness of our approach is demonstrated by simulation studies and an application on predicting health-related outcomes in longitudinal studies of aging.