Data mining, neural nets, trees — Problems 2 and 3 of Genetic Analysis Workshop 15
Data mining, neural nets, trees — Problems 2 and 3 of Genetic Analysis Workshop 15
复制标题
数据挖掘、神经网络、树——遗传分析研讨会 15 的问题 2 和 3
DOI:
10.1002/gepi.20280
复制
发表时间:
2007
影响因子:
2.1
通讯作者:
Zhenming Zhao
中科院分区:
文献类型:
--
作者:
A. Ziegler;A. Destefano;I. König;C. Bardel;D. Brinza;S. Bull;Zhaohui J. Cai;B. Glaser;Wei Jiang;Kristine E. Lee;C. Li;Jing Li;Xin Li;Paul Majoram;Yan Meng;K. Nicodemus;Alexander Platt;D. F. Schwarz;Weilang Shi;Y. Shugart;H. Stassen;Yan V. Sun;Sungho Won;Wenyi Wang;G. Wahba;Usumah A Zagaar;Zhenming Zhao
Genome‐wide association studies using thousands to hundreds of thousands of single nucleotide polymorphism (SNP) markers and region‐wide association studies using a dense panel of SNPs are already in use to identify disease susceptibility genes and to predict disease risk in individuals. Because these tasks become increasingly important, three different data sets were provided for the Genetic Analysis Workshop 15, thus allowing examination of various novel and existing data mining methods for both classification and identification of disease susceptibility genes, gene by gene or gene by environment interaction. The approach most often applied in this presentation group was random forests because of its simplicity, elegance, and robustness. It was used for prediction and for screening for interesting SNPs in a first step. The logistic tree with unbiased selection approach appeared to be an interesting alternative to efficiently select interesting SNPs. Machine learning, specifically ensemble methods, might be useful as pre‐screening tools for large‐scale association studies because they can be less prone to overfitting, can be less computer processor time intensive, can easily include pair‐wise and higher‐order interactions compared with standard statistical approaches and can also have a high capability for classification. However, improved implementations that are able to deal with hundreds of thousands of SNPs at a time are required. Genet. Epidemiol. 31(Suppl. 1):S51–S60, 2007. © 2007 Wiley‐Liss, Inc.
DOI:
10.1007/3-540-45014-9
发表时间:
2000-06
期刊:
--
影响因子:
--
作者:
Thomas G. Dietterich
通讯作者:
Thomas G. Dietterich
影响因子:
3.3
作者:
Templeton,AR;Sing,CF
通讯作者:
Sing,CF