Semi-parametric genomic-enabled prediction of genetic values using reproducing kernel Hilbert spaces methods

Semi-parametric genomic-enabled prediction of genetic values using reproducing kernel Hilbert spaces methods
复制标题

DOI:
10.1017/s0016672310000285
复制
发表时间:
2010-08-01
期刊:
影响因子:
1.5
通讯作者:
Crossa, Jose
Crossa, Jose
中科院分区:
生物学4区
文献类型:
--
作者:
de los Campos, Gustavo;Gianola, Daniel;Crossa, Jose

文献摘要

被引文献

相似文献

遗传值的预测是数量遗传学的核心问题。几十年来,这种预测已经成功地完成了使用表型记录和家庭结构的信息,通常与系谱。现在,人类、植物和动物的基因组中都有密集的分子标记,这些信息可用于加强对遗传价值的预测。然而,将密集的分子标记数据纳入模型提出了许多统计和计算挑战,例如模型如何科普多因子性状的遗传复杂性以及当标记数量超过数据点数量时出现的维数灾难。再生核希尔伯特空间回归可以用来解决其中的一些挑战。该方法允许对几乎任何类型的预测集(协变量,图形,字符串,图像等)进行回归。并且相对于许多参数方法具有重要的计算优势。此外,一些参数模型表现为特殊情况。本文提供了一个概述的方法,讨论了核心选择的问题,重点是遗传应用,算法的核心选择和评估所提出的方法,使用一个集合的599个小麦品系的粮食产量在四个大型环境。
Prediction of genetic values is a central problem in quantitative genetics. Over many decades, such predictions have been successfully accomplished using information on phenotypic records and family structure usually represented with a pedigree. Dense molecular markers are now available in the genome of humans, plants and animals, and this information can be used to enhance the prediction of genetic values. However, the incorporation of dense molecular marker data into models poses many statistical and computational challenges, such as how models can cope with the genetic complexity of multi-factorial traits and with the curse of dimensionality that arises when the number of markers exceeds the number of data points. Reproducing kernel Hilbert spaces regressions can be used to address some of these challenges. The methodology allows regressions on almost any type of prediction sets (covariates, graphs, strings, images, etc.) and has important computational advantages relative to many parametric approaches. Moreover, some parametric models appear as special cases. This article provides an overview of the methodology, a discussion of the problem of kernel choice with a focus on genetic applications, algorithms for kernel selection and an assessment of the proposed methods using a collection of 599 wheat lines evaluated for grain yield in four mega environments.