Variable selection and pattern recognition with gene expression data generated by the microarray technology

Variable selection and pattern recognition with gene expression data generated by the microarray technology
复制标题

DOI:
10.1016/s0025-5564(01)00103-1
复制
发表时间:
2002-03-01
影响因子:
4.3
通讯作者:
Yakovlev, AY
Yakovlev, AY
中科院分区:
生物学4区
文献类型:
--
作者:
Szabo, A;Boucher, K;Yakovlev, AY

文献摘要

被引文献

相似文献

缺乏足够的微阵列数据分析统计方法仍然是揭示这些有前途的技术在基础和转化生物学研究中的真正潜力的最关键的障碍。不应鼓励仅从一张复制品(幻灯片)中得出重要生物学结论的流行做法。在本文中,我们讨论了微阵列数据统计分析的一些现代趋势,特别关注统计分类(模式识别)和变量选择。在解决这些问题时,我们考虑随机向量之间的某些距离及其从基因表达数据获得的非参数估计值的效用。通过计算机模拟和对两种不同类型人类白血病的基因表达数据的分析来测试所提出的距离的性能。在实验设置中,错误率是通过交叉验证来估计的,而控制样本是在计算机模拟实验中生成的,旨在测试所提出的基因选择程序和相关的分类规则。 (C) 2002 Elsevier Science Inc. 保留所有权利。
Lack of adequate statistical methods for the analysis of microarray data remains the most critical deterrent to uncovering the true potential of these promising techniques in basic and translational biological studies. The popular practice of drawing important biological conclusions from just one replicate (slide) should be discouraged. In this paper, we discuss some modern trends in statistical analysis of microarray data with a special focus on statistical classification (pattern recognition) and variable selection. In addressing these issues we consider the utility of some distances between random vectors and their nonparametric estimates obtained from gene expression data. Performance of the proposed distances is tested by computer simulations and analysis of gene expression data on two different types of human leukemia. In experimental settings, the error rate is estimated by cross-validation, while a control sample is generated in computer simulation experiments aimed at testing the proposed gene selection procedures and associated classification rules. (C) 2002 Elsevier Science Inc. All rights reserved.