Evaluation of gene importance in microarray data based upon probability of selection.

Evaluation of gene importance in microarray data based upon probability of selection.
复制标题

DOI:
10.1186/1471-2105-6-67
复制
发表时间:
2005-03-22
期刊:
影响因子:
3
通讯作者:
Fu-Liu CS
Fu-Liu CS
中科院分区:
生物学4区
文献类型:
--
作者:
Fu LM;Fu-Liu CS

文献摘要

参考文献

被引文献

相似文献

微阵列装置允许基因组规模的基因功能评估。近年来,这项技术促进了生物医学的研究和发展。由于许多重要的疾病可以追溯到基因水平,因此一个长期存在的研究问题是确定与导致疾病发展和进展的代谢特征相关的特定基因表达模式。微阵列方法为这个问题提供了一个快速的解决方案。然而,它提出了一个具有挑战性的问题,以识别嵌入在微阵列数据中的疾病相关基因的表达模式。在为分类器设计选择一小组生物学显著基因时,该问题固有的高数据维度的性质产生了大量的不确定性。在这里,我们提出了一个模型的概率分析选定的基因,以确定其重要性。我们的贡献是,我们展示了如何获得的P值的每个选定的基因在多个基因选择试验的数据样本的不同组合的基础上,以及如何进行相应的可靠性分析。基因的重要性由其相关的P值表示,因为从信息论来看,较小的值意味着较高的信息含量。在关于小圆蓝细胞肿瘤亚型分类的微阵列数据上,我们证明了该方法能够找到具有最佳分类性能的最小基因集(19个基因),与文献报道的结果相比。在基于微阵列数据的分类器设计中,基于数据样本的多个组合从基因选择导出的概率值使得能够实现用于减少拟合局部数据特殊性的趋势的有效机制。
Microarray devices permit a genome-scale evaluation of gene function. This technology has catalyzed biomedical research and development in recent years. As many important diseases can be traced down to the gene level, a long-standing research problem is to identify specific gene expression patterns linking to metabolic characteristics that contribute to disease development and progression. The microarray approach offers an expedited solution to this problem. However, it has posed a challenging issue to recognize disease-related genes expression patterns embedded in the microarray data. In selecting a small set of biologically significant genes for classifier design, the nature of high data dimensionality inherent in this problem creates substantial amount of uncertainty. Here we present a model for probability analysis of selected genes in order to determine their importance. Our contribution is that we show how to derive the P value of each selected gene in multiple gene selection trials based on different combinations of data samples and how to conduct a reliability analysis accordingly. The importance of a gene is indicated by its associated P value in that a smaller value implies higher information content from information theory. On the microarray data concerning the subtype classification of small round blue cell tumors, we demonstrate that the method is capable of finding the smallest set of genes (19 genes) with optimal classification performance, compared with results reported in the literature. In classifier design based on microarray data, the probability value derived from gene selection based on multiple combinations of data samples enables an effective mechanism for reducing the tendency of fitting local data particularities.
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V
DOI: 10.1016/s0014-5793(03)00819-6
发表时间: 2003-09-11
期刊: FEBS LETTERS
影响因子: 3.5
作者:
Cho, JH;Lee, D;Lee, IB
通讯作者: Lee, IB
DOI: 10.1109/titb.2003.816558
发表时间: 2003-09-01
影响因子: --
作者:
Fu, LM;Youn, ES
通讯作者: Youn, ES