Data mining, neural nets, trees — Problems 2 and 3 of Genetic Analysis Workshop 15

Data mining, neural nets, trees — Problems 2 and 3 of Genetic Analysis Workshop 15
复制标题

数据挖掘、神经网络、树——遗传分析研讨会 15 的问题 2 和 3

DOI:
10.1002/gepi.20280
复制
发表时间:
2007
影响因子:
2.1
通讯作者:
Zhenming Zhao
Zhenming Zhao
中科院分区:
医学4区
文献类型:
--
作者:
A. Ziegler;A. Destefano;I. König;C. Bardel;D. Brinza;S. Bull;Zhaohui J. Cai;B. Glaser;Wei Jiang;Kristine E. Lee;C. Li;Jing Li;Xin Li;Paul Majoram;Yan Meng;K. Nicodemus;Alexander Platt;D. F. Schwarz;Weilang Shi;Y. Shugart;H. Stassen;Yan V. Sun;Sungho Won;Wenyi Wang;G. Wahba;Usumah A Zagaar;Zhenming Zhao

文献摘要

参考文献

被引文献

相似文献

使用数千到数十万个单核苷酸多态性(SNP)标记的全基因组关联研究和使用密集SNP面板的全区域关联研究已经用于识别疾病易感基因和预测个体的疾病风险。由于这些任务变得越来越重要,因此为遗传分析讲习班15提供了三种不同的数据集,从而允许检查各种新的和现有的数据挖掘方法,以分类和识别疾病易感基因,基因对基因或基因对环境的相互作用。在这个演示组中最常应用的方法是随机森林,因为它简单、优雅和健壮。它在第一步被用于预测和筛选有趣的snp。具有无偏选择方法的逻辑树似乎是有效选择感兴趣snp的一种有趣的替代方法。机器学习,特别是集成方法,可能是用于大规模关联研究的预筛选工具,因为它们不容易出现过拟合,计算机处理器时间密集,与标准统计方法相比,可以很容易地包括成对和高阶相互作用,并且还具有很高的分类能力。但是,需要能够一次处理数十万个snp的改进实现。麝猫。论文,31(增刊。1): S51-S60, 2007年。©2007 Wiley‐Liss, Inc。
Genome‐wide association studies using thousands to hundreds of thousands of single nucleotide polymorphism (SNP) markers and region‐wide association studies using a dense panel of SNPs are already in use to identify disease susceptibility genes and to predict disease risk in individuals. Because these tasks become increasingly important, three different data sets were provided for the Genetic Analysis Workshop 15, thus allowing examination of various novel and existing data mining methods for both classification and identification of disease susceptibility genes, gene by gene or gene by environment interaction. The approach most often applied in this presentation group was random forests because of its simplicity, elegance, and robustness. It was used for prediction and for screening for interesting SNPs in a first step. The logistic tree with unbiased selection approach appeared to be an interesting alternative to efficiently select interesting SNPs. Machine learning, specifically ensemble methods, might be useful as pre‐screening tools for large‐scale association studies because they can be less prone to overfitting, can be less computer processor time intensive, can easily include pair‐wise and higher‐order interactions compared with standard statistical approaches and can also have a high capability for classification. However, improved implementations that are able to deal with hundreds of thousands of SNPs at a time are required. Genet. Epidemiol. 31(Suppl. 1):S51–S60, 2007. © 2007 Wiley‐Liss, Inc.
DOI: 10.1007/3-540-45014-9
发表时间: 2000-06
期刊: --
影响因子: --
作者:
Thomas G. Dietterich
通讯作者: Thomas G. Dietterich
DOI: 10.1093/genetics/134.2.659
发表时间: 1993
期刊: Genetics
影响因子: 3.3
作者:
Templeton,AR;Sing,CF
通讯作者: Sing,CF