Genomic Prediction of Breeding Values Using a Subset of SNPs Identified by Three Machine Learning Methods.

Genomic Prediction of Breeding Values Using a Subset of SNPs Identified by Three Machine Learning Methods.
复制标题

DOI:
10.3389/fgene.2018.00237
复制
发表时间:
2018
影响因子:
3.7
通讯作者:
Li Y
Li Y
中科院分区:
生物学3区
文献类型:
--
作者:
Li B;Zhang N;Wang YG;George AW;Reverter A;Li Y

文献摘要

参考文献

被引文献

相似文献

大型基因组数据的分析受到诸如少量观测和大量预测变量(通常称为“大P小N”)、高维或高度相关的数据结构等问题的阻碍。机器学习方法以处理这些问题而闻名。迄今为止,机器学习方法已应用于全基因组关联研究,用于候选基因的识别,上位性检测,基因网络途径分析和表型值的基因组预测。然而,两种机器学习方法,梯度提升机(GBM)和极端梯度提升方法(XgBoost),在确定一个子集的SNP标记的基因组预测育种值的效用从来没有被探索过。本研究利用2,093头婆罗门牛的38,082个SNP标记和体重表型,(1,097头公牛作为发现群体,996头奶牛作为验证群体),我们检查了三种机器学习方法的效率,即随机森林(RF),GBM和XgBoost,(a)识别排名前400,1,000和3,000的SNP;(B)使用SNP的子集来构建用于估计基因组育种值(GEBV)的基因组关系矩阵(GRM)。出于比较的目的,我们还计算了(1)400,1,000和3,000个随机选择并均匀分布在基因组中的SNP,以及(2)所有SNP的GEBV。我们发现,RF,特别是GBM是有效的方法,在确定一个子集的SNPs与候选基因影响生长性状的直接联系。与使用所有SNP的GEBV预测准确性估计值(0.43)相比,RF(0.42)和GBM(0.46)鉴定的3,000个顶级SNP与整个SNP组的值相似。来自RF和GBM的SNP子集的性能显著优于跨基因组的均匀间隔的子集(0.18-0.29)。在这三种方法中,RF和GBM在基因组预测准确性方面始终优于XgBoost。
The analysis of large genomic data is hampered by issues such as a small number of observations and a large number of predictive variables (commonly known as “large P small N”), high dimensionality or highly correlated data structures. Machine learning methods are renowned for dealing with these problems. To date machine learning methods have been applied in Genome-Wide Association Studies for identification of candidate genes, epistasis detection, gene network pathway analyses and genomic prediction of phenotypic values. However, the utility of two machine learning methods, Gradient Boosting Machine (GBM) and Extreme Gradient Boosting Method (XgBoost), in identifying a subset of SNP makers for genomic prediction of breeding values has never been explored before. In this study, using 38,082 SNP markers and body weight phenotypes from 2,093 Brahman cattle (1,097 bulls as a discovery population and 996 cows as a validation population), we examined the efficiency of three machine learning methods, namely Random Forests (RF), GBM and XgBoost, in (a) the identification of top 400, 1,000, and 3,000 ranked SNPs; (b) using the subsets of SNPs to construct genomic relationship matrices (GRMs) for the estimation of genomic breeding values (GEBVs). For comparison purposes, we also calculated the GEBVs from (1) 400, 1,000, and 3,000 SNPs that were randomly selected and evenly spaced across the genome, and (2) from all the SNPs. We found that RF and especially GBM are efficient methods in identifying a subset of SNPs with direct links to candidate genes affecting the growth trait. In comparison to the estimate of prediction accuracy of GEBVs from using all SNPs (0.43), the 3,000 top SNPs identified by RF (0.42) and GBM (0.46) had similar values to those of the whole SNP panel. The performance of the subsets of SNPs from RF and GBM was substantially better than that of evenly spaced subsets across the genome (0.18–0.29). Of the three methods, RF and GBM consistently outperformed the XgBoost in genomic prediction accuracy.
DOI: 10.1016/j.ygeno.2012.04.003
发表时间: 2012-06
期刊: GENOMICS
影响因子: 4.4
作者:
Chen, Xi;Ishwaran, Hemant
通讯作者: Ishwaran, Hemant
DOI: 10.1016/j.livsci.2014.05.036
发表时间: 2014-08-01
期刊: LIVESTOCK SCIENCE
影响因子: 1.8
作者:
Gonzalez-Recio, Oscar;Rosa, Guilherme J. M.;Gianola, Daniel
通讯作者: Gianola, Daniel
DOI: 10.1534/genetics.108.100289
发表时间: 2009-05-01
期刊: GENETICS
影响因子: 3.3
作者:
Habier, D.;Fernando, R. L.;Dekkers, J. C. M.
通讯作者: Dekkers, J. C. M.
DOI: 10.1534/genetics.112.143313
发表时间: 2013-02
期刊: Genetics
影响因子: 3.3
作者:
de Los Campos G;Hickey JM;Pong-Wong R;Daetwyler HD;Calus MP
通讯作者: Calus MP
DOI: 10.1007/s10709-008-9308-0
发表时间: 2009-06-01
期刊: GENETICA
影响因子: 1.5
作者:
Goddard, Mike
通讯作者: Goddard, Mike