Risk prediction with imperfect survival outcome information from electronic health records.

Risk prediction with imperfect survival outcome information from electronic health records.
复制标题

DOI:
10.1111/biom.13599
复制
发表时间:
2023-03
期刊:
影响因子:
1.9
通讯作者:
Cai, Tianxi
Cai, Tianxi
中科院分区:
数学3区
文献类型:
--
作者:
Hou, Jue;Chan, Stephanie F.;Wang, Xuan;Cai, Tianxi

文献摘要

参考文献

被引文献

相似文献

如果根据较差的代用品进行分析,则容易获得的疾病发病时间代用品,例如第一个诊断代码的时间,可能会导致相当大的风险预测误差。由于缺乏详细的文件和人工注释的劳动强度,通常只能通过随访时间而不是确切时间来确定一小部分疾病的当前状态。在本文中,我们的目标是有效地利用当前状态的少量标签和关于不完美代理的大量未标记观测来建立起病时间的风险预测模型。在半参数变换模型和高度灵活的代理开始时间测量误差模型下,我们提出了一种半监督风险预测方法,该方法将来自代理的信息和来自初始估计器的有限标签有效地结合在一起,仅基于标记的子集,我们用完整的数据对从代理得到的平均零等级相关分数进行一步修正。我们建立了所提出的半监督估计的相合性和渐近正态性质,并给出了区间估计的重采样过程。仿真研究表明,所提出的估计器在有限样本下具有良好的性能。我们通过使用麻省总医院Brigham Healthcare Biobank的数据开发肥胖的遗传风险预测模型来说明所提出的估计器。
Readily available proxies for the time of disease onset such as the time of the first diagnostic code can lead to substantial risk prediction error if performing analyses based on poor proxies. Due to the lack of detailed documentation and labor intensiveness of manual annotation, it is often only feasible to ascertain for a small subset the current status of the disease by a follow-up time rather than the exact time. In this paper, we aim to develop risk prediction models for the onset time efficiently leveraging both a small number of labels on the current status and a large number of unlabeled observations on imperfect proxies. Under a semiparametric transformation model for onset and a highly flexible measurement error model for proxy onset time, we propose the semisupervised risk prediction method by combining information from proxies and limited labels efficiently From an initially estimator solely based on the labeled subset, we perform a one-step correction with the full data augmenting against a mean zero rank correlation score derived from the proxies. We establish the consistency and asymptotic normality of the proposed semisupervised estimator and provide a resampling procedure for interval estimation. Simulation studies demonstrate that the proposed estimator performs well in a finite sample. We illustrate the proposed estimator by developing a genetic risk prediction model for obesity using data from Mass General Brigham Healthcare Biobank.
DOI: 10.1038/ng.686
发表时间: 2010-11
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1002/humu.20822
发表时间: 2009-01
期刊: HUMAN MUTATION
影响因子: 3.9
作者:
Kosoy, Roman;Nassir, Rami;Tian, Chao;White, Phoebe A.;Butler, Lesley M.;Silva, Gabriel;Kittles, Rick;Alarcon-Riquelme, Marta E.;Gregersen, Peter K.;Belmont, John W.;De La Vega, Francisco M.;Seldin, Michael F.
通讯作者: Seldin, Michael F.
DOI: 10.1002/sim.2427
发表时间: 2005-12-30
影响因子: 2
作者:
Antolini, L;Boracchi, P;Biganzoli, E
通讯作者: Biganzoli, E
DOI: 10.1093/biomet/88.2.381
发表时间: 2001-06-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Jin, ZZ;Ying, ZL;Wei, LJ
通讯作者: Wei, LJ
DOI: 10.1016/0304-4076(87)90030-3
发表时间: 1987-07-01
影响因子: 6.3
作者:
HAN, AK
通讯作者: HAN, AK