Genomic Selection in Plant Breeding: A Comparison of Models

Genomic Selection in Plant Breeding: A Comparison of Models
复制标题

DOI:
10.2135/cropsci2011.06.0297
复制
发表时间:
2012-01-01
期刊:
影响因子:
2.3
通讯作者:
Jannink, Jean-Luc
Jannink, Jean-Luc
中科院分区:
农林科学2区
文献类型:
--
作者:
Heslot, Nicolas;Yang, Hsiao-Pei;Jannink, Jean-Luc

文献摘要

被引文献

相似文献

基因组选择(GS)的模拟和实证研究表明,其准确性足以产生快速的遗传增益。然而,随着GS方法的日益普及,已经提出了许多模型,但没有比较分析可以确定最有希望的模型。利用小麦(Triticum aestivum L.)、大麦(Hordeum vulgare L.)、拟南芥(Arabidopsis thaliana L.)Heynh。通过比较每种模型的准确性、基因组估计育种值(gebv)和标记效应,对目前可用的GS模型以及几种机器学习方法的预测能力进行了评估。虽然在许多模型中都观察到类似的精度水平,但过度拟合的水平变化很大,计算时间和标记效应估计的分布也是如此。我们的比较表明,植物育种计划中的GS可以基于一组简化的模型,如贝叶斯套索、加权贝叶斯收缩回归(wBSR, BayesB的快速版本)和随机森林(RF)(一种可以捕获非加性效应的机器学习方法)。测试了不同模型的线性组合以及套袋和增强方法,但它们都没有提高准确性。该研究还表明,数据集中不同亚群之间的准确性差异很大,这并不总是可以用表型方差和大小的差异来解释。这里测试的经验数据集的广泛多样性增加了GS可以增加单位时间和成本的遗传增益的证据。
Simulation and empirical studies of genomic selection (GS) show accuracies sufficient to generate rapid genetic gains. However, with the increased popularity of GS approaches, numerous models have been proposed and no comparative analysis is available to identify the most promising ones. Using eight wheat (Triticum aestivum L.), barley (Hordeum vulgare L.), Arabidopsis thaliana (L.) Heynh., and maize (Zea mays L.) datasets, the predictive ability of currently available GS models along with several machine learning methods was evaluated by comparing accuracies, the genomic estimated breeding values (GEBVs), and the marker effects for each model. While a similar level of accuracy was observed for many models, the level of overfitting varied widely as did the computation time and the distribution of marker effect estimates. Our comparisons suggested that GS in plant breeding programs could be based on a reduced set of models such as the Bayesian Lasso, weighted Bayesian shrinkage regression (wBSR, a fast version of BayesB), and random forest (RF) (a machine learning method that could capture nonadditive effects). Linear combinations of different models were tested as well as bagging and boosting methods, but they did not improve accuracy. This study also showed large differences in accuracy between subpopulations within a dataset that could not always be explained by differences in phenotypic variance and size. The broad diversity of empirical datasets tested here adds evidence that GS could increase genetic gain per unit of time and cost.