SNP selection in genome-wide and candidate gene studies via penalized logistic regression.

SNP selection in genome-wide and candidate gene studies via penalized logistic regression.
复制标题

DOI:
10.1002/gepi.20543
复制
发表时间:
2010-12
影响因子:
2.1
通讯作者:
Cordell, Heather J.
Cordell, Heather J.
中科院分区:
医学4区
文献类型:
--
作者:
Ayers, Kristin L.;Cordell, Heather J.

文献摘要

参考文献

被引文献

相似文献

惩罚回归方法为遗传关联分析中的单标记测试提供了一种有吸引力的替代方法。惩罚回归方法将对感兴趣的性状几乎没有明显影响的标记系数缩小到零,从而产生我们希望的真正相关预测因子的简约子集。在这里,我们探讨了在遗传关联研究中选择 SNP 作为预测因子时惩罚的表现。可以选择惩罚的强度来选择良好的预测模型(通过计算成本高昂的交叉验证等方法),通过基于最大似然的模型选择标准(例如 BIC),或者选择控制 I 类错误的模型,如此处所示。我们研究了几种惩罚逻辑回归方法的性能,模拟了各种疾病位点效应大小和连锁不平衡模式下的数据。我们将 Hyperlasso 软件中实施的几种惩罚方法(包括弹性网、岭、Lasso、MCP 和先验的正态指数 γ 收缩)与标准单轨迹分析和简单的前向逐步回归进行了比较。我们研究了标记如何随着惩罚和 P 值阈值的变化而进入模型,并报告每种方法的敏感性和特异性。结果表明,惩罚方法优于单个标记分析,主要区别在于惩罚方法允许同时包含多个标记,并且通常不允许相关变量进入模型,从而产生稀疏模型,其中考虑了大多数已识别的解释标记。热内特.流行病。 34:879–891, 2010。© 2010 Wiley-Liss, Inc.
Penalized regression methods offer an attractive alternative to single marker testing in genetic association analysis. Penalized regression methods shrink down to zero the coefficient of markers that have little apparent effect on the trait of interest, resulting in a parsimonious subset of what we hope are true pertinent predictors. Here we explore the performance of penalization in selecting SNPs as predictors in genetic association studies. The strength of the penalty can be chosen either to select a good predictive model (via methods such as computationally expensive cross validation), through maximum likelihood-based model selection criterion (such as the BIC), or to select a model that controls for type I error, as done here. We have investigated the performance of several penalized logistic regression approaches, simulating data under a variety of disease locus effect size and linkage disequilibrium patterns. We compared several penalties, including the elastic net, ridge, Lasso, MCP and the normal-exponential-γ shrinkage prior implemented in the hyperlasso software, to standard single locus analysis and simple forward stepwise regression. We examined how markers enter the model as penalties and P-value thresholds are varied, and report the sensitivity and specificity of each of the methods. Results show that penalized methods outperform single marker analysis, with the main difference being that penalized methods allow the simultaneous inclusion of a number of markers, and generally do not allow correlated variables to enter the model, producing a sparse model in which most of the identified explanatory markers are accounted for. Genet. Epidemiol. 34:879–891, 2010. © 2010 Wiley-Liss, Inc.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1371/journal.pgen.1000130
发表时间: 2008-07-25
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Hoggart, Clive J.;Whittaker, John C.;De Iorio, Maria;Balding, David J.
通讯作者: Balding, David J.
DOI: 10.1126/science.1141634
发表时间: 2007-05-11
期刊: SCIENCE
影响因子: 56.9
作者:
Frayling, Timothy M.;Timpson, Nicholas J.;McCarthy, Mark I.
通讯作者: McCarthy, Mark I.