Using penalized regression to predict phenotype from SNP data.

Using penalized regression to predict phenotype from SNP data.
复制标题

DOI:
10.1186/s12919-018-0149-2
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Cordell HJ
Cordell HJ
中科院分区:
其他
文献类型:
--
作者:
Cherlin S;Howey RAJ;Cordell HJ

文献摘要

被引文献

相似文献

在一个典型的基因组预测问题中,预测变量比响应变量多得多。这就禁止了多元线性回归的应用,因为回归系数的唯一普通最小二乘估计量没有定义。为了克服这个问题,已经提出了惩罚回归方法,旨在将系数缩小到零。我们使用惩罚回归方法(LASSO [最小绝对收缩和选择算子]回归)从GAW20数据集中的单核苷酸多态性(SNP)数据探索表型预测。我们使用10倍交叉验证来评估预测性能,并使用10倍嵌套交叉验证来指定惩罚参数。通过分析大约600,000个SNP,我们发现,当样本量包括几百个个体时,SNP效应受到严重惩罚,导致预测性能差。将样本量增加到几千个个体会导致对真实效应的惩罚小得多,从而大大提高了预测。LASSO回归导致回归系数的严重收缩,并且还需要大的样本量(几千个个体)来实现良好的预测。
In a typical genome-enabled prediction problem there are many more predictor variables than response variables. This prohibits the application of multiple linear regression, because the unique ordinary least squares estimators of the regression coefficients are not defined. To overcome this problem, penalized regression methods have been proposed, aiming at shrinking the coefficients toward zero. We explore prediction of phenotype from single nucleotide polymorphism (SNP) data in the GAW20 data set using a penalized regression approach (LASSO [least absolute shrinkage and selection operator] regression). We use 10-fold cross-validation to assess predictive performance and 10-fold nested cross-validation to specify a penalty parameter. By analyzing approximately 600,000 SNPs we find that, when the sample size comprises a few hundred individuals, SNP effects are heavily penalized, resulting in a poor predictive performance. Increasing the sample size to a few thousand individuals results in a much smaller penalization of the true effects, thus greatly improving the prediction. LASSO regression results in a heavy shrinkage of the regression coefficients, and also requires large sample sizes (several thousand individuals) to achieve good prediction.