Dimension reduction for integrative survival analysis

Dimension reduction for integrative survival analysis
复制标题

DOI:
10.1111/biom.13736
复制
发表时间:
2021-08
期刊:
影响因子:
1.9
通讯作者:
Aaron J. Molstad;Rohit Patra
Aaron J. Molstad;Rohit Patra
中科院分区:
数学3区
文献类型:
--
作者:
Aaron J. Molstad;Rohit Patra

文献摘要

相似文献

我们提出了一个约束最大部分似然估计的降维综合(例如,泛癌症)生存分析与高维预测。我们假设对于研究中的每个人群,风险函数遵循不同的考克斯比例风险模型。为了在总体之间借用信息,我们假设每个风险函数仅依赖于少量预测因子的线性组合(即,“因素”)。我们使用基于“距离集”惩罚的算法来估计这些线性组合。这允许我们对回归系数矩阵估计量施加低秩和稀疏性。我们得出的渐近结果表明,我们的估计是更有效的比拟合一个单独的比例风险模型为每个人口。数值实验表明,我们的方法优于竞争对手在各种数据生成模型。我们使用我们的方法进行泛癌症生存分析,将蛋白质表达与18种不同癌症类型的生存率联系起来。我们的方法确定了六种线性组合,仅取决于20种蛋白质,这些蛋白质解释了癌症类型的生存率。最后,为了验证我们的拟合模型,我们证明了我们的估计因子可以在四个外部数据集上比竞争对手更好地预测。
We propose a constrained maximum partial likelihood estimator for dimension reduction in integrative (e.g., pan‐cancer) survival analysis with high‐dimensional predictors. We assume that for each population in the study, the hazard function follows a distinct Cox proportional hazards model. To borrow information across populations, we assume that each of the hazard functions depend only on a small number of linear combinations of the predictors (i.e., “factors”). We estimate these linear combinations using an algorithm based on “distance‐to‐set” penalties. This allows us to impose both low‐rankness and sparsity on the regression coefficient matrix estimator. We derive asymptotic results that reveal that our estimator is more efficient than fitting a separate proportional hazards model for each population. Numerical experiments suggest that our method outperforms competitors under various data generating models. We use our method to perform a pan‐cancer survival analysis relating protein expression to survival across 18 distinct cancer types. Our approach identifies six linear combinations, depending on only 20 proteins, which explain survival across the cancer types. Finally, to validate our fitted model, we show that our estimated factors can lead to better prediction than competitors on four external datasets.