Will Big Data Close the Missing Heritability Gap?

Will Big Data Close the Missing Heritability Gap?
复制标题

DOI:
10.1534/genetics.117.300271
复制
发表时间:
2017-11-01
期刊:
影响因子:
3.3
通讯作者:
de los Campos, Gustavo
de los Campos, Gustavo
中科院分区:
生物学2区
文献类型:
--
作者:
Kim, Hwasoon;Grueneberg, Alexander;de los Campos, Gustavo

文献摘要

被引文献

相似文献

尽管全基因组关联(GWA)研究报告了一些重要发现,但对于大多数性状和疾病,预测的R平方(R-sq.)用遗传分数获得的遗传分数仍然大大低于性状遗传力。现代生物库将很快提供史无前例的大型生物医学数据集:大数据的出现是否会缩小性状遗传性与可由基因组预测因子解释的变异比例之间的差距?我们使用贝叶斯方法和数据分析方法解决了这个问题,该方法产生了与预测R-SQ相关的表面响应。样本大小和模型复杂性(例如,SNP的数量)。我们将这种方法应用于英国生物库中期发布的数据。以身高为模型性状,利用8万条记录进行模型训练,得到了预测R-SQ。在测试中(n=22,221)为0.24(95%C.I.:0.23-0.25)。我们的估计表明,预测的R-SQ。随着样本量的增加,分别使用500个和50,000个(GWA选择的)SNPs的模型达到了估计的平台值,范围从0.1到0.37。很快,更大的数据集就会出现。利用估计的地表响应,我们预测更大的样本量将导致预测R-SQ的进一步改进。我们的结论是,大数据将导致性状遗传力和可用基因组预测因子解释的个体间差异比例之间的差距大幅缩小。然而,即使有了大数据的力量,对于复杂的性状,我们预计预测R-SQ之间的差距。性状遗传性不会完全封闭。
Despite the important discoveries reported by genome-wide association (GWA) studies, for most traits and diseases the prediction R-squared (R-sq.) achieved with genetic scores remains considerably lower than the trait heritability. Modern biobanks will soon deliver unprecedentedly large biomedical data sets: Will the advent of big data close the gap between the trait heritability and the proportion of variance that can be explained by a genomic predictor? We addressed this question using Bayesian methods and a data analysis approach that produces a surface response relating prediction R-sq. with sample size and model complexity (e.g., number of SNPs). We applied the methodology to data from the interim release of the UK Biobank. Focusing on human height as a model trait and using 80,000 records for model training, we achieved a prediction R-sq. in testing (n = 22,221) of 0.24 (95% C. I.: 0.23-0.25). Our estimates show that prediction R-sq. increases with sample size, reaching an estimated plateau at values that ranged from 0.1 to 0.37 for models using 500 and 50,000 (GWA-selected) SNPs, respectively. Soon much larger data sets will become available. Using the estimated surface response, we forecast that larger sample sizes will lead to further improvements in prediction R-sq. We conclude that big data will lead to a substantial reduction of the gap between trait heritability and the proportion of interindividual differences that can be explained with a genomic predictor. However, even with the power of big data, for complex traits we anticipate that the gap between prediction R-sq. and trait heritability will not be fully closed.