A screening methodology based on Random Forests to improve the detection of gene-gene interactions

A screening methodology based on Random Forests to improve the detection of gene-gene interactions
复制标题

DOI:
10.1038/ejhg.2010.48
复制
发表时间:
2010-10-01
影响因子:
5.2
通讯作者:
Van Steen, Kristel
Van Steen, Kristel
中科院分区:
生物学2区
文献类型:
--
作者:
De Lobel, Lizzy;Geurts, Pierre;Van Steen, Kristel

文献摘要

被引文献

相似文献

由于基因-基因相互作用或上位性模型固有的大维度,在基因-基因相互作用中寻找易感位点对统计学家提出了方法和计算上的挑战。在全基因组扫描已经变得相对普遍的时代,需要新的强大方法来处理大量可行的基因-基因相互作用,并从这些结果中剔除假阳性和假阴性。维度问题的一个解决方案是通过初步筛选标记来减少数据,以选择最佳候选者用于进一步分析。理想情况下,该筛选步骤在统计学上独立于测试阶段。最初开发用于少量标记物,多因素简化(MDR)方法是一种非参数、无模型的数据简化技术,用于将标记物集与疾病的最佳预测特性相关联。在这项研究中,我们在更大的数据集中研究了MDR的能力,并将其与其他能够识别基因-基因相互作用的方法进行了比较。在各种交互模型(纯上位和非纯上位)下,我们使用基于随机森林(RF)的预筛选方法,在执行MDR之前,以提高其性能。我们发现,MDR的功率增加时,噪声SNP首先被删除,通过创建一个集合的候选标记与RF。我们通过广泛的模拟研究和欧洲呼吸健康研究II委员会哮喘数据的应用验证了我们的技术。European Journal of Human Genetics(2010)18,1127-1132; doi:10.1038/ejhg.2010.48; 2010年5月12日在线发表
The search for susceptibility loci in gene-gene interactions imposes a methodological and computational challenge for statisticians because of the large dimensionality inherent to the modelling of gene-gene interactions or epistasis. In an era in which genome-wide scans have become relatively common, new powerful methods are required to handle the huge amount of feasible gene-gene interactions and to weed out false positives and negatives from these results. One solution to the dimensionality problem is to reduce data by preliminary screening of markers to select the best candidates for further analysis. Ideally, this screening step is statistically independent of the testing phase. Initially developed for small numbers of markers, the Multifactor Dimensionality Reduction (MDR) method is a nonparametric, model-free data reduction technique to associate sets of markers with optimal predictive properties to disease. In this study, we examine the power of MDR in larger data sets and compare it with other approaches that are able to identify gene-gene interactions. Under various interaction models (purely and not purely epistatic), we use a Random Forest (RF)-based prescreening method, before executing MDR, to improve its performance. We find that the power of MDR increases when noisy SNPs are first removed, by creating a collection of candidate markers with RFs. We validate our technique by extensive simulation studies and by application to asthma data from the European Committee of Respiratory Health Study II. European Journal of Human Genetics (2010) 18, 1127-1132; doi: 10.1038/ejhg.2010.48; published online 12 May 2010