Tumor classification by partial least squares using microarray gene expression data

Tumor classification by partial least squares using microarray gene expression data
复制标题

DOI:
10.1093/bioinformatics/18.1.39
复制
发表时间:
2002-01-01
期刊:
影响因子:
5.8
通讯作者:
Rocke, DM
Rocke, DM
中科院分区:
生物学3区
文献类型:
--
作者:
Nguyen, DV;Rocke, DM

文献摘要

被引文献

相似文献

动机:基因表达微阵列数据的一个重要应用是将样本分类,如肿瘤的类型。微阵列的使用允许同时监测每个样本的数千个基因表达。这种整体测量基因表达的能力导致了变量p(基因)的数量远远超过样本N的数量的数据。当N<p;p时,标准的统计分类和预测方法不能很好地甚至根本不起作用。为了分析微阵列数据,需要修改现有的统计方法或开发新的方法。结果:我们提出了一种基于微阵列基因表达的人类肿瘤样本分类(预测)的新分析方法。这一过程包括使用偏最小二乘(PLS)降维以及使用Logistic判别(LD)和二次判别分析(QDA)进行分类。将最小二乘法与主成分分析的降维方法进行了比较。在许多情况下,偏最小二乘被证明是优越的;我们举例说明了主成分分析相对于偏最小二乘法尤其不能很好地预测的情况。所提出的方法被应用于涉及不同人类肿瘤样本的五个不同的微阵列数据集:(1)正常与卵巢肿瘤;(2)急性髓系白血病(AML)与急性淋巴细胞白血病(ALL);(3)差异使用大B细胞淋巴瘤(DLBCLL)与B细胞慢性淋巴细胞性白血病(BCLL);(4)正常与结肠癌;以及(5)非小细胞肺癌(NSCLC)与肾脏样本。通过再随机化研究进一步评估分类结果和方法的稳定性。
Motivation: One important application of gene expression microarray data is classification of samples into categories, such as the type of tumor. The use of microarrays allows simultaneous monitoring of thousands of genes expressions per sample. This ability to measure gene expression en masse has resulted in data with the number of variables p (genes) far exceeding the number of samples N. Standard statistical methodologies in classification and prediction do not work well or even at all when N < p. Modification of existing statistical methodologies or development of new methodologies is needed for the analysis of microarray data.Results: We propose a novel analysis procedure for classifying (predicting) human tumor samples based on microarray gene expressions. This procedure involves dimension reduction using Partial Least Squares (PLS) and classification using Logistic Discrimination (LD) and Quadratic Discriminant Analysis (QDA). We compare PLS to the well known dimension reduction method of Principal Components Analysis (PCA). Under many circumstances PLS proves superior; we illustrate a condition when PCA particularly fails to predict well relative to PLS. The proposed methods were applied to five different microarray data sets involving various human tumor samples: (1) normal versus ovarian tumor; (2) Acute Myeloid Leukemia (AML) versus Acute Lymphoblastic Leukemia (ALL); (3) Diff use Large B-cell Lymphoma (DLBCLL) versus B-cell Chronic Lymphocytic Leukemia (BCLL); (4) normal versus colon tumor; and (5) Non-Small-Cell-Lung-Carcinoma (NSCLC) versus renal samples. Stability of classification results and methods were further assessed by re-randomization studies.