Efficient strategies for leave-one-out cross validation for genomic best linear unbiased prediction.

Efficient strategies for leave-one-out cross validation for genomic best linear unbiased prediction.
复制标题

DOI:
10.1186/s40104-017-0164-6
复制
发表时间:
2017
影响因子:
7
通讯作者:
Fernando RL
Fernando RL
中科院分区:
农林科学1区
文献类型:
--
作者:
Cheng H;Garrick DJ;Fernando RL

文献摘要

被引文献

相似文献

利用全基因组数据,提出了一种随机多元回归模型,该模型将加性标记或单倍型的所有等位基因替代效应同时拟合为不相关的随机效应,用于最佳线性无偏预测。留一法交叉验证可用于量化统计模型的预测能力。留一交叉验证的简单应用是计算密集型的,因为训练和验证分析需要重复n次,每次观察一次。这里提出了有效的留一交叉验证策略,只需要比单个分析多一点的努力。高效的留一法交叉验证策略对于具有1,000个观测值和10,000个标记的模拟数据集比朴素应用程序快786倍,对于具有1,000个观测值和100个标记的模拟数据集快99倍。相对于使用相同模型的朴素方法,这些效率将随着观测数量的增加而增加。这里提出了有效的留一交叉验证策略,只需要比单个分析多一点的努力。
A random multiple-regression model that simultaneously fit all allele substitution effects for additive markers or haplotypes as uncorrelated random effects was proposed for Best Linear Unbiased Prediction, using whole-genome data. Leave-one-out cross validation can be used to quantify the predictive ability of a statistical model. Naive application of Leave-one-out cross validation is computationally intensive because the training and validation analyses need to be repeated n times, once for each observation. Efficient Leave-one-out cross validation strategies are presented here, requiring little more effort than a single analysis. Efficient Leave-one-out cross validation strategies is 786 times faster than the naive application for a simulated dataset with 1,000 observations and 10,000 markers and 99 times faster with 1,000 observations and 100 markers. These efficiencies relative to the naive approach using the same model will increase with increases in the number of observations. Efficient Leave-one-out cross validation strategies are presented here, requiring little more effort than a single analysis.