Performance of random forest when SNPs are in linkage disequilibrium

Performance of random forest when SNPs are in linkage disequilibrium
复制标题

DOI:
10.1186/1471-2105-10-78
复制
发表时间:
2009-03-05
期刊:
影响因子:
3
通讯作者:
Lunetta, Kathryn L.
Lunetta, Kathryn L.
中科院分区:
生物学4区
文献类型:
--
作者:
Meng, Yan A.;Yu, Yi;Lunetta, Kathryn L.

文献摘要

被引文献

相似文献

背景:单核苷酸多态性(SNP)可能由于连锁不平衡(LD)而相关。关联研究寻找与疾病位点的直接和间接关联。在随机森林(RF)分析中,真实风险SNP与LD中的SNP之间的相关性可能导致真实风险SNP的变量重要性降低。解决这个问题的一种方法是选择连锁平衡(LE)中的SNP进行分析。在这里,我们探讨了其他方法来处理单核苷酸多态性在LD:改变树的建设算法,通过建立每个树在RF中只与单核苷酸多态性在LE,修改的重要性措施(IM),并使用单倍型,而不是单核苷酸多态性,以建立一个RF.Results:我们评估了我们的替代方法的性能,通过模拟的频谱复杂的遗传模型。当单倍型而不是单个SNP是风险因素时,我们发现对SNP进行的原始随机森林方法提供了良好的性能。当个体、基因型SNP是风险因素时,我们发现遗传效应越强,LD对原始RF性能的影响越强。与原始RF一起使用的修订的重要性度量在SNP中对LD相对稳健;与修订的RF一起使用的该修订的重要性度量有时被夸大。总体而言,我们发现,当遗传模型和LD中具有风险SNP的SNP数量未知时,与原始RF一起使用的修订的重要性度量是最佳选择。对于单倍型为基础的方法,在乘法异质性模型下,我们观察到RF的性能下降,增加LD之间的SNPs在haplotype.Conclusion:我们的研究结果表明,通过战略性地修改随机森林方法树建设或重要性度量计算,功率可以增加时,LD之间存在的SNPs。我们的结论是,修改后的随机森林方法进行SNP提供了一个优势,不需要基因型阶段,使其成为一个可行的工具,用于在成千上万的SNP的背景下,如候选基因研究和后续的顶级候选人从全基因组关联研究。
Background: Single nucleotide polymorphisms (SNPs) may be correlated due to linkage disequilibrium (LD). Association studies look for both direct and indirect associations with disease loci. In a Random Forest (RF) analysis, correlation between a true risk SNP and SNPs in LD may lead to diminished variable importance for the true risk SNP. One approach to address this problem is to select SNPs in linkage equilibrium (LE) for analysis. Here, we explore alternative methods for dealing with SNPs in LD: change the tree-building algorithm by building each tree in an RF only with SNPs in LE, modify the importance measure (IM), and use haplotypes instead of SNPs to build a RF.Results: We evaluated the performance of our alternative methods by simulation of a spectrum of complex genetics models. When a haplotype rather than an individual SNP is the risk factor, we find that the original Random Forest method performed on SNPs provides good performance. When individual, genotyped SNPs are the risk factors, we find that the stronger the genetic effect, the stronger the effect LD has on the performance of the original RF. A revised importance measure used with the original RF is relatively robust to LD among SNPs; this revised importance measure used with the revised RF is sometimes inflated. Overall, we find that the revised importance measure used with the original RF is the best choice when the genetic model and the number of SNPs in LD with risk SNPs are unknown. For the haplotype-based method, under a multiplicative heterogeneity model, we observed a decrease in the performance of RF with increasing LD among the SNPs in the haplotype.Conclusion: Our results suggest that by strategically revising the Random Forest method tree-building or importance measure calculation, power can increase when LD exists between SNPs. We conclude that the revised Random Forest method performed on SNPs offers an advantage of not requiring genotype phase, making it a viable tool for use in the context of thousands of SNPs, such as candidate gene studies and follow-up of top candidates from genome wide association studies.