Partial least squares regression, support vector machine regression, and transcriptome-based distances for prediction of maize hybrid performance with gene expression data

Partial least squares regression, support vector machine regression, and transcriptome-based distances for prediction of maize hybrid performance with gene expression data
复制标题

DOI:
10.1007/s00122-011-1747-9
复制
发表时间:
2012-03-01
影响因子:
5.4
通讯作者:
Frisch, Matthias
Frisch, Matthias
中科院分区:
农林科学1区
文献类型:
--
作者:
Fu, Junjie;Falke, K. Christin;Frisch, Matthias

文献摘要

被引文献

相似文献

杂交种的性能可以用来自其亲本近交系的基因表达数据来预测。在育种计划中实施这种预测方法有望提高杂交育种的效率。我们研究的目的是比较预测模型的准确性,采用多元线性回归(MLR),偏最小二乘回归(PLS),支持向量机回归(SVM),和基于转录组的距离(DB)。对于7个燧石玉米系和14个马齿玉米系的阶乘,评估杂交种的谷粒产量,并用56 k微阵列分析亲本系的基因表达。预测模型的准确性进行了衡量的预测和观察到的产量采用两个交叉验证计划之间的相关性。第一个模型预测的杂交种时,测交数据可用于两个亲本系(2型杂交种),第二个模型预测的杂交种时,没有测交数据的亲本系(0型杂交种)。MLR、SVM和PLS对2型杂交种的预测产量和观测产量之间的相关性较高,而对0型杂交种D-B的预测精度较高。回归方法对分析基因集的选择是稳健的,并且仅需要几百个基因。相比之下,用D-B进行精确的杂交预测,需要1,000 - 1,500个基因,并且预测准确性强烈依赖于所分析的基因的集合。我们的结论是,在一组遗传物质的预测MLR是一个很有前途的方法,并从一组遗传物质的预测模型转移到相关的,基于转录组的距离D-B是最有前途的。
The performance of hybrids can be predicted with gene expression data from their parental inbred lines. Implementing such prediction approaches in breeding programs promises to increase the efficiency of hybrid breeding. The objectives of our study were to compare the accuracy of prediction models employing multiple linear regression (MLR), partial least squares regression (PLS), support vector machine regression (SVM), and transcriptome-based distances (DB). For a factorial of 7 flint and 14 dent maize lines, the grain yield of the hybrids was assessed and the gene expression of the parental lines was profiled with a 56k microarray. The accuracy of the prediction models was measured by the correlation between predicted and observed yield employing two cross-validation schemes. The first modeled the prediction of hybrids when testcross data are available for both parental lines (type 2 hybrids), and the second modeled the prediction of hybrids when no testcross data for the parental lines were available (type 0 hybrids). MLR, SVM, and PLS resulted in a high correlation between predicted and observed yield for type 2 hybrids, whereas for type 0 hybrids D-B had greater prediction accuracy. The regression methods were robust to the choice of the set of profiled genes and required only a few hundred genes. In contrast, for an accurate hybrid prediction with D-B, 1,000-1,500 genes were required, and the prediction accuracy depended strongly on the set of profiled genes. We conclude that for prediction within one set of genetic material MLR is a promising approach, and for transfering prediction models from one set of genetic material to a related one, the transcriptome-based distance D-B is most promising.