Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanks.

Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanks.
复制标题

使用带时间戳的电子健康记录数据进行胰腺癌表型风险评分 (PheRS):在两个大型生物库中的发现和验证。

DOI:
10.1016/j.jbi.2020.103652
复制
发表时间:
2021-01
影响因子:
4.5
通讯作者:
Mukherjee B
Mukherjee B
中科院分区:
医学3区
文献类型:
--
作者:
Salvatore M;Beesley LJ;Fritsche LG;Hanauer D;Shi X;Mondul AM;Pearce CL;Mukherjee B

文献摘要

参考文献

被引文献

相似文献

用于疾病风险预测和评估的传统方法,例如使用血清、尿液、血液、唾液或成像生物标志物的诊断测试,对于识别许多疾病的高风险个体,导致早期检测和改善存活率非常重要。对于胰腺癌,传统的筛查方法在疾病进展之前识别高风险个体方面基本上是不成功的,导致高死亡率和低生存率。与基因图谱相关的电子健康记录(EHR)为整合多个来源的患者信息进行风险预测和分层提供了机会。我们利用EHR中可用的时间相关诊断的星座来构建一个汇总风险评分,称为表型风险评分(PheRS),用于识别患有胰腺癌的高风险个体。拟议的PheRS方法将疾病发作的时间纳入预测框架。我们联合收割机将PheRS与更知名的遗传易感性指标,即预测胰腺癌的多基因风险评分(PRS)相结合并进行对比。我们首先计算了胰腺癌诊断与医学表型组中所有可能的其他诊断之间的成对、未调整的关联。我们称这些成对关联为同现。在考虑交叉表型相关性后,使用来自相对独立诊断子集的多变量关联估计值来创建加权和PheRS。我们使用来自密歇根基因组计划(MGI)的38,359名参与者的数据,根据EHR中包含的在目标胰腺癌诊断前0年,1年,2年和5年的诊断构建了时间限制的风险评分。在英国生物样本库(UKB)中评估PheRS的可预测性。我们测试了PheRS在添加到包含遗传易感性(PRS)汇总指标加上其他协变量(如年龄、性别、吸烟状况、饮酒状况和体重指数(BMI))的模型中时的相对贡献。我们对共现模式的探索确定了预期的关联,同时也揭示了可能需要密切关注的意外关系。在UKB中,仅使用目标诊断前5年的胰腺癌PheRS得出的AUC为0.60(95% CI = [0.58,0.62])。在UKB中,包括PheRS、PRS和协变量的5年阈值的更大预测模型的AUC为0.74(95% CI = [0.72,0.76])。我们注意到,PheRS在联合模型中确实有独立的贡献。最后,在PheRS分布的顶部的分数证明了在风险分层方面的前景。在胰腺癌诊断前5年阈值时,UKB中前2%的评分识别病例的可能性是后98%的评分的10.20倍(95% CI = [9.34,12.99])。我们开发了一个框架,用于使用医学表型组的丰富信息内容从胰腺癌的EHR数据创建有时间限制的PheRS。除了为未来的研究确定产生假设的关联之外,即使在调整了胰腺癌和其他传统流行病学协变量的PRS之后,该PheRS也在识别高风险个体方面表现出潜在的重要贡献。该方法可推广到其他表型性状。
Traditional methods for disease risk prediction and assessment, such as diagnostic tests using serum, urine, blood, saliva or imaging biomarkers, have been important for identifying high-risk individuals for many diseases, leading to early detection and improved survival. For pancreatic cancer, traditional methods for screening have been largely unsuccessful in identifying high-risk individuals in advance of disease progression leading to high mortality and poor survival. Electronic health records (EHR) linked to genetic profiles provide an opportunity to integrate multiple sources of patient information for risk prediction and stratification. We leverage a constellation of temporally associated diagnoses available in the EHR to construct a summary risk score, called a phenotype risk score (PheRS), for identifying individuals at high-risk for having pancreatic cancer. The proposed PheRS approach incorporates the time with respect to disease onset into the prediction framework. We combine and contrast the PheRS with more well-known measures of inherited susceptibility, namely, the polygenic risk scores (PRS) for prediction of pancreatic cancer. We first calculated pairwise, unadjusted associations between pancreatic cancer diagnosis and all possible other diagnoses across the medical phenome. We call these pairwise associations co-occurrences. After accounting for cross-phenotype correlations, the multivariable association estimates from a subset of relatively independent diagnoses were used to create a weighted sum PheRS. We constructed time-restricted risk scores using data from 38,359 participants in the Michigan Genomics Initiative (MGI) based on the diagnoses contained in the EHR at 0, 1, 2, and 5 years prior to the target pancreatic cancer diagnosis. The PheRS was assessed for predictability in the UK Biobank (UKB). We tested the relative contribution of PheRS when added to a model containing a summary measure of inherited genetic susceptibility (PRS) plus other covariates like age, sex, smoking status, drinking status, and body mass index (BMI). Our exploration of co-occurrence patterns identified expected associations while also revealing unexpected relationships that may warrant closer attention. Solely using the pancreatic cancer PheRS at 5 years before the target diagnoses yielded an AUC of 0.60 (95% CI = [0.58, 0.62]) in UKB. A larger predictive model including PheRS, PRS, and the covariates at the 5-year threshold achieved an AUC of 0.74 (95% CI = [0.72, 0.76]) in UKB. We note that PheRS does contribute independently in the joint model. Finally, scores at the top percentiles of the PheRS distribution demonstrated promise in terms of risk stratification. Scores in the top 2% were 10.20 (95% CI = [9.34, 12.99]) times more likely to identify cases than those in the bottom 98% in UKB at the 5-year threshold prior to pancreatic cancer diagnosis. We developed a framework for creating a time-restricted PheRS from EHR data for pancreatic cancer using the rich information content of a medical phenome. In addition to identifying hypothesis-generating associations for future research, this PheRS demonstrates a potentially important contribution in identifying high-risk individuals, even after adjusting for PRS for pancreatic cancer and other traditional epidemiologic covariates. The methods are generalizable to other phenotypic traits.
DOI: 10.1038/nbt.2749
发表时间: 2013-12
影响因子: 46.9
作者:
通讯作者: --
DOI: 10.1126/science.aal4043
发表时间: 2018-03-16
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Bastarache L;Hughey JJ;Hebbring S;Marlo J;Zhao W;Ho WT;Van Driest SL;McGregor TL;Mosley JD;Wells QS;Temple M;Ramirez AH;Carroll R;Osterman T;Edwards T;Ruderfer D;Velez Edwards DR;Hamid R;Cogan J;Glazer A;Wei WQ;Feng Q;Brilliant M;Zhao ZJ;Cox NJ;Roden DM;Denny JC
通讯作者: Denny JC
DOI: 10.1093/nar/gky1120
发表时间: 2019-01-08
影响因子: 14.9
作者:
Buniello, Annalisa;MacArthur, Jacqueline A. L.;Parkinson, Helen
通讯作者: Parkinson, Helen
DOI: 10.1093/bioinformatics/btu197
发表时间: 2014-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Carroll, Robert J.;Bastarache, Lisa;Denny, Joshua C.
通讯作者: Denny, Joshua C.
DOI: 10.1016/j.ajhg.2018.04.001
发表时间: 2018-06-07
影响因子: 9.8
作者:
Fritsche, Lars G.;Gruber, Stephen B.;Mukherjee, Bhramar
通讯作者: Mukherjee, Bhramar