Machine learning methods and predictive ability metrics for genome-wide prediction of complex traits

Machine learning methods and predictive ability metrics for genome-wide prediction of complex traits
复制标题

DOI:
10.1016/j.livsci.2014.05.036
复制
发表时间:
2014-08-01
期刊:
影响因子:
1.8
通讯作者:
Gianola, Daniel
Gianola, Daniel
中科院分区:
农林科学3区
文献类型:
--
作者:
Gonzalez-Recio, Oscar;Rosa, Guilherme J. M.;Gianola, Daniel

文献摘要

被引文献

相似文献

复杂性状的全基因组预测在动植物育种中变得越来越重要,并在人类遗传学中受到越来越多的关注。最常见的方法是全基因组回归模型,其中同时回归数千个标记的表型,将不同的先验分布应用于标记效应。虽然在SNP回归模型中使用收缩或正则化提高了基于基因组的评估的预测能力,但随着标记和可用表型之间的比率持续增加,可能会遇到严重的过度拟合问题。机器学习是预测和分类的一种替代方法,能够以计算灵活的方式处理维度问题。在本文中,我们提供了非参数和机器学习方法用于全基因组预测的概述,讨论了它们的相似性以及它们与一些著名的参数方法的关系。虽然最适合的方法通常是依赖于案例,但我们建议使用支持向量机和随机森林来解决分类问题,而再生核Hilbert空间回归和Boosting可能适合更好的回归问题,前者具有更一致的更高的预测能力。神经网络在神经元数目较多时可能存在过度拟合和计算量过大的问题,本文从基因组选择的角度进一步讨论了交叉验证下模型比较中用于评估预测能力的指标。我们建议使用预测均方误差作为模型比较的主要指标,但不仅仅是衡量标准。可视化工具可以极大地帮助选择最准确的模型。(C)2014爱思唯尔B.V.保留所有权利。
Genome-wide prediction of complex traits has become increasingly important in animal and plant breeding, and is receiving increasing attention in human genetics. Most common approaches are whole-genome regression models where phenotypes are regressed on thousands of markers concurrently, applying different prior distributions to marker effects. While use of shrinkage or regularization in SNP regression models has delivered improvements in predictive ability in genome-based evaluations, serious over-fitting problems may be encountered as the ratio between markers and available phenotypes continues increasing. Machine learning is an alternative approach for prediction and classification, capable of dealing with the dimensionality problem in a computationally flexible manner. In this article we provide an overview of non-parametric and machine learning methods used in genome wide prediction, discuss their similarities as well as their relationship to some well-known parametric approaches. Although the most suitable method is usually case dependent, we suggest the use of support vector machines and random forests for classification problems, whereas Reproducing Kernel Hilbert Spaces regression and boosting may suit better regression problems, with the former having the more consistently higher predictive ability. Neural Networks may suffer from over-fitting and may be too computationally demanded when the number of neurons is large.We further discuss on the metrics used to evaluate predictive ability in model comparison under cross-validation from a genomic selection point of view. We suggest use of predictive mean squared error as a main but not only metric for model comparison. Visual tools may greatly assist on the choice of the most accurate model. (C) 2014 Elsevier B.V. All rights reserved.