Multi-class cancer classification via partial least squares with gene expression profiles

Multi-class cancer classification via partial least squares with gene expression profiles
复制标题

DOI:
10.1093/bioinformatics/18.9.1216
复制
发表时间:
2002-09-01
期刊:
影响因子:
5.8
通讯作者:
Rocke, DM
Rocke, DM
中科院分区:
生物学3区
文献类型:
--
作者:
Nguyen, DV;Rocke, DM

文献摘要

被引文献

相似文献

动机:基于基因表达谱的两类样品(例如正常样品和癌症样品)之间以及两种类型的癌症之间的区分是一个重要的问题,其具有实际意义以及进一步我们理解各种癌细胞的基因表达的潜力。还需要对两个以上的群体或类别(多类别)进行分类或区分。多类歧视的方法的需要是显而易见的,在许多微阵列实验中,不同的癌症类型被认为是simultaneous.Results:因此,在本文中,我们提出了早期的分类方法的扩展阮和洛克(2002年b;生物信息学,18,39-50),从多个类别的癌症样本进行分类。本文中提出的方法被应用于具有多个类别的四个基因表达数据集:(a)遗传性乳腺癌数据集,其具有(1)BRCA 1-突变、(2)BRCA 2-突变和(3)散发性乳腺癌样品,(B)急性白血病数据集,其具有(1)急性髓性白血病(AML),(2)T细胞急性淋巴细胞白血病(T-ALL)和(3)B细胞急性淋巴细胞白血病(B-ALL)样品,(c)淋巴瘤数据集,具有(1)弥漫性大B细胞淋巴瘤(DLBCL),(2)B细胞慢性淋巴细胞白血病(BCLL)和(3)滤泡性淋巴瘤(FL)样品,和(d)具有源自各种起源部位的癌症的细胞系的NCI 60数据集。此外,我们评估了分类算法,并使用基于真实的数据集随机化的模拟来检查错误率的变化。我们注意到,最近有其他方法用于解决多类预测,我们的方法是沿着Nguyen和Rocke(2002 b; Bioinformatics,18,39-50)的路线。
Motivation: Discrimination between two classes such as normal and cancer samples and between two types of cancers based on gene expression profiles is an important problem which has practical implications as well as the potential to further our understanding of gene expression of various cancer cells. Classification or discrimination of more than two groups or classes (multi-class) is also needed. The need for multi-class discrimination methodologies is apparent in many microarray experiments where various cancer types are considered simultaneously.Results: Thus, in this paper we present the extension to the classification methodology proposed earlier Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50) to classify cancer samples from multiple classes. The methodologies proposed in this paper are applied to four gene expression data sets with multiple classes: (a) a hereditary breast cancer data set with (1) BRCA1-mutation, (2) BRCA2-mutation and (3) sporadic breast cancer samples, (b) an acute leukemia data set with (1) acute myeloid leukemia (AML), (2) T-cell acute lymphoblastic leukemia (T-ALL) and (3) B-cell acute lymphoblastic leukemia (B-ALL) samples, (c) a lymphoma data set with (1) diffuse large B-cell lymphoma (DLBCL), (2) B-cell chronic lymphocytic leukemia (BCLL) and (3) follicular lymphoma (FL) samples, and (d) the NCI60 data set with cell lines derived from cancers of various sites of origin. In addition, we evaluated the classification algorithms and examined the variability of the error rates using simulations based on randomization of the real data sets. We note that there are other methods for addressing multi-class prediction recently and our approach is along the line of Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50).