Benchmarking Parametric and Machine Learning Models for Genomic Prediction of Complex Traits

Benchmarking Parametric and Machine Learning Models for Genomic Prediction of Complex Traits
复制标题

DOI:
10.1534/g3.119.400498
复制
发表时间:
2019-11-01
影响因子:
2.6
通讯作者:
Shiu, Shin-Han
Shiu, Shin-Han
中科院分区:
生物学3区
文献类型:
--
作者:
Azodi, Christina B.;Bolger, Emily;Shiu, Shin-Han

文献摘要

被引文献

相似文献

基因组预测在作物和牲畜育种计划中的有用性促使人们努力开发新的和改进的基因组预测算法,如人工神经网络和梯度树提升。然而,这些算法的性能还没有使用广泛的数据集和模型以系统的方式进行比较。利用不同标记密度和训练群体大小的6个植物物种的18个性状的数据,我们比较了6种线性和6种非线性算法的性能。首先,我们发现超参数选择对于所有非线性算法都是必要的,并且当标记的数量远远超过训练线的数量时,模型训练之前的特征选择对于人工神经网络来说是至关重要的。在所有物种和性状组合中,没有一种算法表现最好,但是基于多个算法的结果组合的预测(即,整体预测)表现一贯良好。虽然线性和非线性算法对相似数量的性状表现最好,但非线性算法的性能在性状之间变化更大。虽然人工神经网络在任何性状上都没有表现得最好,但我们确定了策略(即,特征选择,种子起始权重),将其性能提升到接近其他算法的水平。我们的研究结果突出了算法选择的重要性,性状值的预测。
The usefulness of genomic prediction in crop and livestock breeding programs has prompted efforts to develop new and improved genomic prediction algorithms, such as artificial neural networks and gradient tree boosting. However, the performance of these algorithms has not been compared in a systematic manner using a wide range of datasets and models. Using data of 18 traits across six plant species with different marker densities and training population sizes, we compared the performance of six linear and six non-linear algorithms. First, we found that hyperparameter selection was necessary for all non-linear algorithms and that feature selection prior to model training was critical for artificial neural networks when the markers greatly outnumbered the number of training lines. Across all species and trait combinations, no one algorithm performed best, however predictions based on a combination of results from multiple algorithms (i.e., ensemble predictions) performed consistently well. While linear and non-linear algorithms performed best for a similar number of traits, the performance of non-linear algorithms vary more between traits. Although artificial neural networks did not perform best for any trait, we identified strategies (i.e., feature selection, seeded starting weights) that boosted their performance to near the level of other algorithms. Our results highlight the importance of algorithm selection for the prediction of trait values.