Tissue classification with gene expression profiles

Tissue classification with gene expression profiles
复制标题

DOI:
10.1089/106652700750050943
复制
发表时间:
2000-01-01
影响因子:
1.7
通讯作者:
Yakhini, Z
Yakhini, Z
中科院分区:
生物学4区
文献类型:
--
作者:
Ben-Dor, A;Bruhn, L;Yakhini, Z

文献摘要

被引文献

相似文献

不断改善基因表达分析技术有望提供对与癌症相关的细胞过程的理解和见解。基因表达数据还有望大大有助于发展癌症诊断和分类平台的发展。在这项工作中,我们检查了三组基因表达数据,这些基因表达数据在一组肿瘤和正常临床样本中测量:第一组由2,000个基因组成,在62个上皮结肠样品中测量(Alon等,1999)。第二个由大约100,000个克隆组成,在32个卵巢样品中测量(Schummer等人(1999)中描述的数据集的未发表扩展)。第三组由大约7,100个基因组成,在72个骨髓和周围育雏样品中测量(Golub等,1999)。我们检查了评分方法的使用,测量使用单个基因表达水平的组织类型分离(例如,肿瘤与正常肿瘤)。然后将它们与高维分类方法结合在一起,以评估完整表达曲线的分类能力。我们介绍了在三个数据集上执行一对一的交叉验证(LOOCV)实验的结果,该数据集使用最近的邻居分类器SVM(Cortes and Vapnik,1995),Adaboost(Freund and Schapire,1997)和一个基于新颖的聚类集群分类技术。由于肿瘤样品与其细胞型组成中的正常样品可能有所不同,因此我们还使用适当修改的基因集进行了LOOCV实验,试图消除所得偏差。我们证明,使用一组选定的基因,以及与细胞 - 污染相关的成员,肿瘤与正常分类的成功率至少为90%。这些结果对一定范围内的确切选择机制不敏感。
Constantly improving gene expression profiling technologies are expected to provide understanding and insight into cancer-related cellular processes. Gene expression data is also expected to significantly aid in the development of efficient cancer diagnosis and classification platforms. In this work we examine three sets of gene expression data measured across sets of tumor(s) and normal clinical samples: The first set consists of 2,000 genes, measured in 62 epithelial colon samples (Alon et al., 1999). The second consists of approximate to 100,000 clones, measured in 32 ovarian samples (unpublished extension of data set described in Schummer et al, (1999)). The third set consists of approximate to 7,100 genes, measured in 72 bone marrow and peripheral brood samples (Golub et al., 1999). We examine the use of scoring methods, measuring separation of tissue type (e.g., tumors from normals) using individual gene expression levels. These are then coupled with high-dimensional classification methods to assess the classification power of complete expression profiles. We present results of performing leave-one-one cross validation (LOOCV) experiments on the three data sets, employing nearest neighbor classifier, SVM (Cortes and Vapnik, 1995), AdaBoost (Freund and Schapire, 1997) and a novel clustering-based classification technique. As tumor samples can differ from normal samples in their cell-type composition, we also perform LOOCV experiments using appropriately modified sets of genes, attempting to eliminate the resulting bias. We demonstrate success rate of at least 90% in tumor versus normal classification, using sets of selected genes, with, as well as without, cellular-contamination-related members. These results are insensitive to the exact selection mechanism, over a certain range.