A computationally efficient hypothesis testing method for epistasis analysis using multifactor dimensionality reduction.

A computationally efficient hypothesis testing method for epistasis analysis using multifactor dimensionality reduction.
复制标题

DOI:
10.1002/gepi.20360
复制
发表时间:
2009-01
影响因子:
2.1
通讯作者:
Moore, Jason H.
Moore, Jason H.
中科院分区:
医学4区
文献类型:
--
作者:
Pattin, Kristine A.;White, Bill C.;Barney, Nate;Gui, Jiang;Nelson, Heather H.;Kelsey, Karl T.;Andrew, Angeline S.;Karagas, Margaret R.;Moore, Jason H.

文献摘要

参考文献

被引文献

相似文献

多因素降维(MDR)是一种非参数和无模型的数据挖掘方法,用于检测,表征和解释上位性,在遗传和流行病学研究中缺乏显着的主效应的复杂性状,如疾病易感性。MDR的目标是使用构造性归纳算法改变数据的表示,使非加性相互作用更容易使用任何分类方法(如朴素贝叶斯或逻辑回归)检测。传统上,MDR构造变量已经使用朴素贝叶斯分类器进行评估,该分类器与10倍交叉验证相结合,以获得预测准确性或上位性模型的普遍性的估计。传统上,我们使用排列检验来统计评估通过MDR获得的模型的显著性。置换测试的优点是它可以控制由于多次测试而导致的假阳性。缺点是置换测试在计算上是昂贵的。这是在全基因组范围内检测上位性的背景下出现的一个重要问题。本研究的目的是开发和评估几种替代大规模排列检验评估MDR模型的统计学意义。使用70种不同上位性模型模拟的数据,我们使用1000倍排列检验和使用极值分布(EVD)的假设检验比较MDR的功效和I型错误率。我们发现,这种新的假设检验方法提供了一个合理的替代计算昂贵的1000倍排列测试,是50倍的速度。然后,我们通过将其应用于膀胱癌易感性的遗传流行病学研究来证明这种新方法,该研究先前使用MDR进行分析并使用1000倍排列检验进行评估。
Multifactor dimensionality reduction (MDR) was developed as a nonparametric and model-free data mining method for detecting, characterizing, and interpreting epistasis in the absence of significant main effects in genetic and epidemiologic studies of complex traits such as disease susceptibility. The goal of MDR is to change the representation of the data using a constructive induction algorithm to make nonadditive interactions easier to detect using any classification method such as naïve Bayes or logistic regression. Traditionally, MDR constructed variables have been evaluated with a naïve Bayes classifier that is combined with 10-fold cross validation to obtain an estimate of predictive accuracy or generalizability of epistasis models. Traditionally, we have used permutation testing to statistically evaluate the significance of models obtained through MDR. The advantage of permutation testing is that it controls for false-positives due to multiple testing. The disadvantage is that permutation testing is computationally expensive. This is in an important issue that arises in the context of detecting epistasis on a genome-wide scale. The goal of the present study was to develop and evaluate several alternatives to large-scale permutation testing for assessing the statistical significance of MDR models. Using data simulated from 70 different epistasis models, we compared the power and type I error rate of MDR using a 1000-fold permutation test with hypothesis testing using an extreme value distribution (EVD). We find that this new hypothesis testing method provides a reasonable alternative to the computationally expensive 1000-fold permutation test and is 50 times faster. We then demonstrate this new method by applying it to a genetic epidemiology study of bladder cancer susceptibility that was previously analyzed using MDR and assessed using a 1000-fold permutation test.
DOI: 10.1159/000073735
发表时间: 2003-01-01
期刊: HUMAN HEREDITY
影响因子: 1.8
作者:
Moore, JH
通讯作者: Moore, JH
DOI: 10.1086/321276
发表时间: 2001-07-01
影响因子: 9.8
作者:
Ritchie, MD;Hahn, LW;Moore, JH
通讯作者: Moore, JH
DOI: 10.1086/423738
发表时间: 2004-09-01
影响因子: 9.8
作者:
Dudbridge, F;Koeleman, BPC
通讯作者: Koeleman, BPC
DOI: 10.1007/bf00532484
发表时间: 1983-01-01
期刊: ZEITSCHRIFT FUR WAHRSCHEINLICHKEITSTHEORIE UND VERWANDTE GEBIETE
影响因子: --
作者:
LEADBETTER, MR
通讯作者: LEADBETTER, MR
DOI: 10.1093/bioinformatics/btf869
发表时间: 2003-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hahn, LW;Ritchie, MD;Moore, JH
通讯作者: Moore, JH