Evaluating the Ability of Tree-Based Methods and Logistic Regression for the Detection of SNP-SNP Interaction

Evaluating the Ability of Tree-Based Methods and Logistic Regression for the Detection of SNP-SNP Interaction
复制标题

DOI:
10.1111/j.1469-1809.2009.00511.x
复制
发表时间:
2009-05-01
影响因子:
1.9
通讯作者:
Salas, Antonio
Salas, Antonio
中科院分区:
生物学4区
文献类型:
--
作者:
Garcia-Magarinos, Manuel;Lopez-de-Ullibarri, Inaki;Salas, Antonio

文献摘要

被引文献

相似文献

大多数常见的人类疾病可能有复杂的病因。允许上位现象的分析方法在复杂疾病的遗传解剖中越来越引起人们的兴趣。通过允许潜在疾病位点之间的上位相互作用,我们可能成功地识别出可能未被发现的遗传变异。在这里,我们的目的是分析逻辑回归(LR)和两个基于树的监督学习方法,分类和回归树(CART)和随机森林(RF),检测上位性的能力。多因素降维(MDR)也用于比较。我们的方法首先涉及常染色体双等位基因非定相和非连锁单核苷酸多态性(SNP)的数据集的模拟,每个包含两个位点的相互作用(因果SNP)和98个“噪音”SNP。我们在不同的情况下的样本量,缺失数据,次要等位基因频率(MAF)和几个重复率模型的相互作用建模:三个涉及(无法区分的)边际效应和相互作用,和两个模拟纯相互作用的影响。我们总共模拟了99种不同的场景。虽然CART、RF和LR在检测真实关联方面产生类似的结果,但CART和RF在分类错误方面的表现优于LR。MAF、回归模型和样本大小是不同技术检测真实关联能力的决定性因素,而不是缺失数据百分比。在纯交互模型中,只有RF检测关联。总之,基于树的方法和LR是重要的统计工具,用于检测真实风险相关SNP之间的未知相互作用,具有边际效应,并且存在大量噪声SNP。在纯交互作用模型中,RF在大样本量和低百分比缺失数据的情况下表现相当好。然而,当研究设计不理想时(在样本量和MAF方面不利于检测相互作用),检测到虚假、虚假关联的可能性很高。
Most common human diseases are likely to have complex etiologies. Methods of analysis that allow for the phenomenon of epistasis are of growing interest in the genetic dissection of complex diseases. By allowing for epistatic interactions between potential disease loci, we may succeed in identifying genetic variants that might otherwise have remained undetected. Here we aimed to analyze the ability of logistic regression (LR) and two tree-based supervised learning methods, classification and regression trees (CART) and random forest (RF), to detect epistasis. Multifactor-dimensionality reduction (MDR) was also used for comparison. Our approach involves first the simulation of datasets of autosomal biallelic unphased and unlinked single nucleotide polymorphisms (SNPs), each containing a two-loci interaction (causal SNPs) and 98 'noise' SNPs. We modelled interactions under different scenarios of sample size, missing data, minor allele frequencies (MAF) and several penetrance models: three involving both (indistinguishable) marginal effects and interaction, and two simulating pure interaction effects. In total, we have simulated 99 different scenarios. Although CART, RF, and LR yield similar results in terms of detection of true association, CART and RF perform better than LR with respect to classification error. MAF, penetrance model, and sample size are greater determining factors than percentage of missing data in the ability of the different techniques to detect true association. In pure interaction models, only RF detects association. In conclusion, tree-based methods and LR are important statistical tools for the detection of unknown interactions among true risk-associated SNPs with marginal effects and in the presence of a significant number of noise SNPs. In pure interaction models, RF performs reasonably well in the presence of large sample sizes and low percentages of missing data. However, when the study design is suboptimal (unfavourable to detect interaction in terms of e.g. sample size and MAF) there is a high chance of detecting false, spurious associations.