Components of the accuracy of genomic prediction in a multi-breed sheep population

Components of the accuracy of genomic prediction in a multi-breed sheep population
复制标题

DOI:
10.2527/jas.2011-4557
复制
发表时间:
2012-10-01
影响因子:
3.3
通讯作者:
Hayes, B. J.
Hayes, B. J.
中科院分区:
农林科学2区
文献类型:
--
作者:
Daetwyler, H. D.;Kemper, K. E.;Hayes, B. J.

文献摘要

被引文献

相似文献

在全基因组关联研究中,由于群体结构而未能消除变异导致虚假关联。相比之下,对于从密集SNP数据预测未来表型或估计育种值,利用由相关性产生的群体结构实际上可以在某些情况下增加预测的准确性,例如,当选择候选者是预测方程导出的参考群体的后代时。在具有大有效群体大小或具有多个品种和品系的群体中,尚未证明是否以及何时考虑或消除由于群体结构引起的变异将影响基因组预测的准确性。我们在这项研究中的目的是确定是否占人口结构将增加基因组预测的准确性,无论是在品种内和品种之间。首先,我们试图将基因组预测的准确性分解为来自群体结构或标记与QTL之间的连锁不平衡(LD)的贡献,使用不同的多品种绵羊(Ovis aries)数据集,对48,640个SNP进行基因分型。我们证明,来自单个染色体的SNP可以达到使用所有SNP的基因组预测的准确性的86%。这一结果表明,大部分的预测精度是由于人口结构,因为一个单一的染色体预计捕捉关系,但不太可能包含所有的QTL。然后,我们探讨了主成分分析(PCA)作为一种方法来解开各自的贡献的群体结构和LD之间的SNP和QTL的基因组预测的准确性。结果表明,拟合越来越多的主成分(PC;作为协变量)下降品种内的准确性,直到达到一个较低的平台。我们推测,这个平台是由于LD的准确性的措施。总之,在我们的数据中,基因组预测的准确性很大一部分是由于与种群结构相关的变异。令人惊讶的是,考虑到这种结构通常会降低跨品种基因组预测的准确性。
In genome-wide association studies, failure to remove variation due to population structure results in spurious associations. In contrast, for predictions of future phenotypes or estimated breeding values from dense SNP data, exploiting population structure arising from relatedness can actually increase the accuracy of prediction in some cases, for example, when the selection candidates are offspring of the reference population where the prediction equation was derived. In populations with large effective population size or with multiple breeds and strains, it has not been demonstrated whether and when accounting for or removing variation due to population structure will affect the accuracy of genomic prediction. Our aim in this study was to determine whether accounting for population structure would increase the accuracy of genomic predictions, both within and across breeds. First, we have attempted to decompose the accuracy of genomic prediction into contributions from population structure or linkage disequilibrium (LD) between markers and QTL using a diverse multi-breed sheep (Ovis aries) data set, genotyped for 48,640 SNP. We demonstrate that SNP from a single chromosome can achieve up to 86% of the accuracy for genomic predictions using all SNP. This result suggests that most of the prediction accuracy is due to population structure, because a single chromosome is expected to capture relationships but is unlikely to contain all QTL. We then explored principal component analysis (PCA) as an approach to disentangle the respective contributions of population structure and LD between SNP and QTL to the accuracy of genomic predictions. Results showed that fitting an increasing number of principle components (PC; as covariates) decreased within breed accuracy until a lower plateau was reached. We speculate that this plateau is a measure of the accuracy due to LD. In conclusion, a large proportion of the accuracy for genomic predictions in our data was due to variation associated with population structure. Surprisingly, accounting for this structure generally decreased the accuracy of across breed genomic predictions.