Non-parametric Polygenic Risk Prediction via Partitioned GWAS Summary Statistics

Non-parametric Polygenic Risk Prediction via Partitioned GWAS Summary Statistics
复制标题

DOI:
10.1016/j.ajhg.2020.05.004
复制
发表时间:
2020-07-02
影响因子:
9.8
通讯作者:
Sunyaev, Shamil R.
Sunyaev, Shamil R.
中科院分区:
生物学1区
文献类型:
--
作者:
Chun, Sung;Imakaev, Maxim;Sunyaev, Shamil R.

文献摘要

被引文献

相似文献

在复杂性状遗传学中,从基因型预测表型的能力是我们理解性状遗传力的遗传结构的最终衡量标准。一个完整的理解的遗传基础的性状应允许预测方法的准确性接近性状的遗传力。数量性状和最常见的表型的高度多基因性促使了统计策略的发展,这些策略集中于组合无数个体非显著的遗传效应。现在,预测准确性正在提高,有越来越多的兴趣,在实际效用的这种方法来预测常见疾病的风险响应早期治疗干预。然而,现有的方法需要个体水平的基因型或依赖于准确地指定潜在的每种疾病的遗传结构进行预测。在这里,我们提出了一个多基因的风险预测方法,不需要明确建模任何潜在的遗传结构。我们从一个大型GWAS队列的SNP效应量形式的汇总统计开始。然后,我们删除了由于连锁不平衡而产生的汇总统计量之间的相关性结构,并对条件均值效应应用分段线性插值。在模拟和真实的数据集中,这种新的非参数收缩(NST)方法可以可靠地允许500万个密集的全基因组标记的汇总统计中的连锁不平衡,并持续提高预测精度。我们表明,ESTA提高了乳腺癌,2型糖尿病,炎症性肠病和冠心病的高风险群体的识别,所有这些都有可用的早期干预或预防治疗。
In complex trait genetics, the ability to predict phenotype from genotype is the ultimate measure of our understanding of genetic architecture underlying the heritability of a trait. A complete understanding of the genetic basis of a trait should allow for predictive methods with accuracies approaching the trait's heritability. The highly polygenic nature of quantitative traits and most common phenotypes has motivated the development of statistical strategies focused on combining myriad individually non-significant genetic effects. Now that predictive accuracies are improving, there is a growing interest in the practical utility of such methods for predicting risk of common diseases responsive to early therapeutic intervention. However, existing methods require individual-level genotypes or depend on accurately specifying the genetic architecture underlying each disease to be predicted. Here, we propose a polygenic risk prediction method that does not require explicitly modeling any underlying genetic architecture. We start with summary statistics in the form of SNP effect sizes from a large GWAS cohort. We then remove the correlation structure across summary statistics arising due to linkage disequilibrium and apply a piecewise linear interpolation on conditional mean effects. In both simulated and real datasets, this new non-parametric shrinkage (NPS) method can reliably allow for linkage disequilibrium in summary statistics of 5 million dense genome-wide markers and consistently improves prediction accuracy. We show that NPS improves the identification of groups at high risk for breast cancer, type 2 diabetes, inflammatory bowel disease, and coronary heart disease, all of which have available early intervention or prevention treatments.