Accuracy of breeding values of 'unrelated' individuals predicted by dense SNP genotyping

Accuracy of breeding values of 'unrelated' individuals predicted by dense SNP genotyping
复制标题

DOI:
10.1186/1297-9686-41-35
复制
发表时间:
2009-06-11
影响因子:
4.1
通讯作者:
Meuwissen, Theo H. E.
Meuwissen, Theo H. E.
中科院分区:
生物学2区
文献类型:
--
作者:
Meuwissen, Theo H. E.

文献摘要

被引文献

相似文献

背景资料:SNP发现和高通量基因分型技术的最新发展使得使用高密度SNP标记来预测育种值变得可行。这涉及在训练数据集中估计SNP效应,并使用这些估计来评估其他“评估”个体的育种值。模拟研究表明,当训练和评估个体(密切)相关时,这些育种值的预测可以是准确的。然而,基因组选择的许多一般应用需要预测的育种值的“无关”的个人,即来自同一人口的个人,但不是特别密切相关的培训individual.Methods:选择的准确性进行了调查,通过计算机模拟的小群体。使用缩放参数,结果被扩展到不同的人口,训练数据集和基因组大小,和不同的性状heritability.Results:预测育种值的无关个体需要一个显着更高的标记密度和数量的训练记录时,预测个人的后代的训练个人。而当记录数为2*N-e*L,标记数为10*N-e*L时,可预测无关个体的育种值,其准确率为0.88 ~ 0.93,其中Ne为有效群体大小,L为Morgan基因组大小。将这一要求降低到1*N-e*L个体,预测精度降低到0.73- 0.83。结论:对于家畜群体,1 N(e)L需要大约30,000个训练记录,但如果训练和评估动物相关,这一要求可能会降低。提出了一个预测方程,当训练和评估个体相关时,该预测方程预测精度。对于人类来说,1 NeL需要近似35万个个体,这意味着人类疾病风险预测只可能用于由有限数量的基因决定的疾病。否则,基因分型和表型记录需要在未来变得非常普遍。
Background: Recent developments in SNP discovery and high throughput genotyping technology have made the use of high-density SNP markers to predict breeding values feasible. This involves estimation of the SNP effects in a training data set, and use of these estimates to evaluate the breeding values of other 'evaluation' individuals. Simulation studies have shown that these predictions of breeding values can be accurate, when training and evaluation individuals are (closely) related. However, many general applications of genomic selection require the prediction of breeding values of 'unrelated' individuals, i.e. individuals from the same population, but not particularly closely related to the training individuals.Methods: Accuracy of selection was investigated by computer simulation of small populations. Using scaling arguments, the results were extended to different populations, training data sets and genome sizes, and different trait heritabilities.Results: Prediction of breeding values of unrelated individuals required a substantially higher marker density and number of training records than when prediction individuals were offspring of training individuals. However, when the number of records was 2*N-e*L and the number of markers was 10*N-e*L, the breeding values of unrelated individuals could be predicted with accuracies of 0.88 - 0.93, where Ne is the effective population size and L the genome size in Morgan. Reducing this requirement to 1*N-e*L individuals, reduced prediction accuracies to 0.73-0.83.Conclusion: For livestock populations, 1N(e)L requires about similar to 30,000 training records, but this may be reduced if training and evaluation animals are related. A prediction equation is presented, that predicts accuracy when training and evaluation individuals are related. For humans, 1NeL requires similar to 350,000 individuals, which means that human disease risk prediction is possible only for diseases that are determined by a limited number of genes. Otherwise, genotyping and phenotypic recording need to become very common in the future.