Identifying marker genes in transcription profiling data using a mixture of feature relevance experts

Identifying marker genes in transcription profiling data using a mixture of feature relevance experts
复制标题

DOI:
10.1152/physiolgenomics.2001.5.2.99
复制
发表时间:
2001-03-08
影响因子:
4.6
通讯作者:
Mian, IS
Mian, IS
中科院分区:
生物学3区
文献类型:
--
作者:
Chow, ML;Moler, EJ;Mian, IS

文献摘要

被引文献

相似文献

转录图谱实验允许同时测量许多基因的表达水平。给出两种类型样本的特征数据,最能区分样本的基因(标记基因)是后续深入实验研究和开发用于诊断、预后和监测的决策支持系统的良好候选。这项工作提出了一种混合特征相关性专家作为识别标记基因的方法,并使用来自标记为急性淋巴细胞白血病和髓系白血病(ALL,AML)的样本的已发表数据来说明这一想法。特征相关性专家实现了一种算法,该算法计算基因区分样本的好坏,根据该相关性度量对基因重新排序,并使用监督学习方法[这里是支持向量机(SVMs)]来确定不同嵌套基因子集的泛化性能。研究的三个特征相关性专家的混合实施了两个现有的和一个新的特征相关性度量。对于每一位专家,由前50个基因组成的基因子集将所有AML样本与所有7070个基因完全区分开来。排名前50位的125个基因可能是原型决策支持系统的标志物。染色体异常和其他数据支持这样的预测,即位于前50位的三个基因,胱抑素C,天青素和脂蛋白,是研究ALL/AML基础生物学的良好靶点。使用相同的数据来识别基于T细胞/B细胞、外周血/骨髓和男性/女性的标记来区分样本的标记。硒蛋白W可以区分T细胞和B细胞。来自肿瘤/非肿瘤结肠腺癌样本的转录图谱数据的分析结果支持上述方法的一般实用性。讨论了支持向量机核函数及其参数的选择、特征相关专家的训练和评估以及潜在错误标记样本对标记识别(特征选择)的影响等理论问题。
Transcription profiling experiments permit the expression levels of many genes to be measured simultaneously. Given profiling data from two types of samples, genes that most distinguish the samples (marker genes) are good candidates for subsequent in-depth experimental studies and developing decision support systems for diagnosis, prognosis, and monitoring. This work proposes a mixture of feature relevance experts as a method for identifying marker genes and illustrates the idea using published data from samples labeled as acute lymphoblastic and myeloid leukemia (ALL, AML). A feature relevance expert implements an algorithm that calculates how well a gene distinguishes samples, reorders genes according to this relevance measure, and uses a supervised learning method [here, support vector machines (SVMs)] to determine the generalization performances of different nested gene subsets. The mixture of three feature relevance experts examined implement two existing and one novel feature relevance measures. For each expert, a gene subset consisting of the top 50 genes distinguished ALL from AML samples as completely as all 7,070 genes. The 125 genes at the union of the top 50s are plausible markers for a prototype decision support system. Chromosomal aberration and other data support the prediction that the three genes at the intersection of the top 50s, cystatin C, azurocidin, and adipsin, are good targets for investigating the basic biology of ALL/AML. The same data were employed to identify markers that distinguish samples based on their labels of T cell/B cell, peripheral blood/bone marrow, and male/female. Selenoprotein W may discriminate T cells from B cells. Results from analysis of transcription profiling data from tumor/nontumor colon adenocarcinoma samples support the general utility of the aforementioned approach. Theoretical issues such as choosing SVM kernels and their parameters, training and evaluating feature relevance experts, and the impact of potentially mislabeled samples on marker identification (feature selection) are discussed.