Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores

Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores
复制标题

DOI:
10.1016/j.ajhg.2015.09.001
复制
发表时间:
2015-10-01
影响因子:
9.8
通讯作者:
Price, Alkes L.
Price, Alkes L.
中科院分区:
生物学1区
文献类型:
--
作者:
Vilhjalmsson, Bjarni J.;Yang, Jian;Price, Alkes L.

文献摘要

被引文献

相似文献

多基因风险评分在预测复杂疾病风险方面表现出很大的希望,并且随着训练样本量的增加将变得更加准确。计算风险分数的标准方法涉及基于连锁不平衡(LD)的标记修剪和将p值阈值应用于关联统计,但这会丢弃信息并降低预测准确性。我们介绍LDpred,一种通过使用来自外部参考面板的效应大小和LD信息的先验来推断每个标记物的后验平均效应大小的方法。理论和仿真结果表明,LDpred优于修剪阈值,特别是在大样本量的方法。因此,在大型精神分裂症数据集中,预测的R-2从20.1%增加到25.3%,在大型多发性硬化症数据集中,从9.8%增加到12.0%。在另外三个大型疾病数据集和非欧洲精神分裂症样本中观察到类似的准确性相对改善。LDpred相对于现有方法的优势将随着样本量的增加而增加。
Polygenic risk scores have shown great promise in predicting complex disease risk and will become more accurate as training sample sizes increase. The standard approach for calculating risk scores involves linkage disequilibrium (LD)-based marker pruning and applying a p value threshold to association statistics, but this discards information and can reduce predictive accuracy. We introduce LDpred, a method that infers the posterior mean effect size of each marker by using a prior on effect sizes and LD information from an external reference panel. Theory and simulations show that LDpred outperforms the approach of pruning followed by thresholding, particularly at large sample sizes. Accordingly, predicted R-2 increased from 20.1% to 25.3% in a large schizophrenia dataset and from 9.8% to 12.0% in a large multiple sclerosis dataset. A similar relative improvement in accuracy was observed for three additional large disease datasets and for non-European schizophrenia samples. The advantage of LDpred over existing methods will grow as sample sizes increase.