Using Genetic Distance to Infer the Accuracy of Genomic Prediction.

Using Genetic Distance to Infer the Accuracy of Genomic Prediction.
复制标题

DOI:
10.1371/journal.pgen.1006288
复制
发表时间:
2016-09
期刊:
影响因子:
4.5
通讯作者:
Balding D
Balding D
中科院分区:
生物学2区
文献类型:
--
作者:
Scutari M;Mackay I;Balding D

文献摘要

参考文献

被引文献

相似文献

利用高密度基因组数据预测表型性状有许多应用,如具有商业价值的动植物的选择;它有望在医学诊断中发挥越来越大的作用。用于此任务的统计模型通常使用交叉验证进行测试,这隐含地假设新个体(我们想要预测其表型)来自基因组预测模型所训练的同一种群。本文提出了一种基于聚类和重抽样的方法来研究增加训练群体与目标群体之间的遗传距离对数量性状预测的影响。这对植物和动物遗传学很重要,因为基因组选择程序依赖于对未来育种预测的准确性。因此,估计预测精度衰减的速度对于决定使用哪个训练群体以及模型需要重新校准的频率是很重要的。我们发现真实值和预测值之间的相关性与训练和目标群体之间的FST或平均亲属关系近似线性衰减。我们通过模拟和收集来自小鼠、小麦和人类遗传学的数据集来说明这种关系。越来越多的基因组数据的可用性使得使用统计模型来预测感兴趣的特征成为生命科学中许多应用的支柱。应用范围从常见和罕见疾病的医学诊断到具有商业利益的动植物的抗病育种特性。我们探索了一个关于如何评估这种预测模型的隐含假设:我们想要预测其特征的个体与用于训练模型的个体来自同一种群。通常情况并非如此,特别是在作为选择程序一部分的植物和动物的情况下。为了研究这个问题,我们提出了一种模型不可知的方法来推断预测模型的准确性作为两个常见的遗传距离度量的函数。利用植物、动物和人类遗传学的数据,我们发现这两种测量方法的准确性都近似线性衰减。量化这种衰退在遗传学的所有分支中都有基本的应用,因为它可以衡量研究如何推广到不同的人群。
The prediction of phenotypic traits using high-density genomic data has many applications such as the selection of plants and animals of commercial interest; and it is expected to play an increasing role in medical diagnostics. Statistical models used for this task are usually tested using cross-validation, which implicitly assumes that new individuals (whose phenotypes we would like to predict) originate from the same population the genomic prediction model is trained on. In this paper we propose an approach based on clustering and resampling to investigate the effect of increasing genetic distance between training and target populations when predicting quantitative traits. This is important for plant and animal genetics, where genomic selection programs rely on the precision of predictions in future rounds of breeding. Therefore, estimating how quickly predictive accuracy decays is important in deciding which training population to use and how often the model has to be recalibrated. We find that the correlation between true and predicted values decays approximately linearly with respect to either FST or mean kinship between the training and the target populations. We illustrate this relationship using simulations and a collection of data sets from mice, wheat and human genetics. The availability of increasing amounts of genomic data is making the use of statistical models to predict traits of interest a mainstay of many applications in life sciences. Applications range from medical diagnostics for common and rare diseases to breeding characteristics such as disease resistance in plants and animals of commercial interest. We explored an implicit assumption of how such prediction models are often assessed: that the individuals whose traits we would like to predict originate from the same population as those that are used to train the models. This is commonly not the case, especially in the case of plants and animals that are parts of selection programs. To study this problem we proposed a model-agnostic approach to infer the accuracy of prediction models as a function of two common measures of genetic distance. Using data from plant, animal and human genetics, we find that accuracy decays approximately linearly in either of those measures. Quantifying this decay has fundamental applications in all branches of genetics, as it measures how studies generalise to different populations.
DOI: 10.1186/1297-9686-42-5
发表时间: 2010-02-19
期刊: Genetics, selection, evolution : GSE
影响因子: --
作者:
Habier D;Tetens J;Seefried FR;Lichtner P;Thaller G
通讯作者: Thaller G
DOI: 10.1534/genetics.107.081190
发表时间: 2007-12-01
期刊: GENETICS
影响因子: 3.3
作者:
Habier, D.;Fernando, R. L.;Dekkers, J. C. M.
通讯作者: Dekkers, J. C. M.
DOI: 10.2135/cropsci2013.03.0195
发表时间: 2014-07-01
期刊: CROP SCIENCE
影响因子: 2.3
作者:
Hickey, John M.;Dreisigacker, Susanne;Gorjanc, Gregor
通讯作者: Gorjanc, Gregor
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1007/s10709-008-9308-0
发表时间: 2009-06-01
期刊: GENETICA
影响因子: 1.5
作者:
Goddard, Mike
通讯作者: Goddard, Mike