On the statistical assessment of classifiers using DNA microarray data

On the statistical assessment of classifiers using DNA microarray data
复制标题

DOI:
10.1186/1471-2105-7-387
复制
发表时间:
2006-08-19
期刊:
影响因子:
3
通讯作者:
Perri, F.
Perri, F.
中科院分区:
生物学4区
文献类型:
--
作者:
Ancona, N.;Maglietta, R.;Perri, F.

文献摘要

被引文献

相似文献

背景:在本文中,我们提出了一种利用基因表达谱对癌症预测因子进行统计评估的方法。该方法应用于在意大利福贾 Casa Sollievo della Sofferenza 医院收集的微阵列基因表达数据的新数据集。该数据集由从 25 名结肠癌患者中提取的正常样本 (22) 和肿瘤样本 (25) 组成。我们建议回答一些与癌症自动诊断相关的问题,例如:可用数据集的大小是否足以构建准确的分类器?相关错误率的统计显着性是什么?根据所采用的分类方案,可以通过哪些方式来考虑准确性?有多少基因与病理学相关,有多少基因足以进行准确的结肠癌分类?我们提出的方法回答了这些问题,同时避免了微阵列数据分析和解释中隐藏的潜在陷阱。结果:我们通过改变训练示例的数量和使用的基因的数量,通过Leave-K-Out交叉验证误差评估三种不同分类方案的泛化误差。错误率的统计显着性通过使用排列检验来测量。我们提供了分类中涉及的基因频率的统计分析。使用整组基因,我们发现加权投票算法 (WVA) 分类器通过 25 个训练示例学习正常样本和肿瘤样本之间的区别,提供 e = 21% (p = 0.045) 的错误率。即使示例数量增加,这个值仍然保持不变。此外,正则化最小二乘法 (RLS) 和支持向量机 (SVM) 分类器只需 15 个训练样本即可学习,错误率分别为 e = 19% (p = 0.035) 和 e = 18% (p = 0.037)。此外,错误率随着训练集大小的增加而降低,在 35 个训练样本时达到最佳性能。在这种情况下,RLS 和 SVM 的错误率为 e = 14% (p = 0.027) 和 e = 11% (p = 0.019)。关于基因数量,根据信噪比统计,我们发现大约 6000 个基因 (p < 0.05) 与病理相关。此外,当使用 74% 的基因时,RLS 和 SVM 分类器的性能不会改变。当仅使用 2 个基因时,它们逐渐减少至 e = 16% (p < 0.05)。讨论了通过我们的统计分析确定的一组基因的生物学相关性以及它们在结直肠肿瘤发生中发挥的主要作用。结论:所提出的方法为与癌症诊断和预后相关的精确问题提供了统计上显着的答案。我们发现,只需 15 个示例,就可以训练出具有统计显着性的分类器来进行结肠癌诊断。至于足以对结肠癌进行可靠分类的基因数量的定义,我们的结果表明,这取决于所需的准确性。
Background: In this paper we present a method for the statistical assessment of cancer predictors which make use of gene expression profiles. The methodology is applied to a new data set of microarray gene expression data collected in Casa Sollievo della Sofferenza Hospital, Foggia - Italy. The data set is made up of normal (22) and tumor (25) specimens extracted from 25 patients affected by colon cancer. We propose to give answers to some questions which are relevant for the automatic diagnosis of cancer such as: Is the size of the available data set sufficient to build accurate classifiers? What is the statistical significance of the associated error rates? In what ways can accuracy be considered dependant on the adopted classification scheme? How many genes are correlated with the pathology and how many are sufficient for an accurate colon cancer classification? The method we propose answers these questions whilst avoiding the potential pitfalls hidden in the analysis and interpretation of microarray data.Results: We estimate the generalization error, evaluated through the Leave-K-Out Cross Validation error, for three different classification schemes by varying the number of training examples and the number of the genes used. The statistical significance of the error rate is measured by using a permutation test. We provide a statistical analysis in terms of the frequencies of the genes involved in the classification. Using the whole set of genes, we found that the Weighted Voting Algorithm (WVA) classifier learns the distinction between normal and tumor specimens with 25 training examples, providing e = 21% (p = 0.045) as an error rate. This remains constant even when the number of examples increases. Moreover, Regularized Least Squares (RLS) and Support Vector Machines (SVM) classifiers can learn with only 15 training examples, with an error rate of e = 19% (p = 0.035) and e = 18% (p = 0.037) respectively. Moreover, the error rate decreases as the training set size increases, reaching its best performances with 35 training examples. In this case, RLS and SVM have error rates of e = 14% (p = 0.027) and e = 11% (p = 0.019). Concerning the number of genes, we found about 6000 genes (p < 0.05) correlated with the pathology, resulting from the signal-to-noise statistic. Moreover the performances of RLS and SVM classifiers do not change when 74% of genes is used. They progressively reduce up to e = 16% (p < 0.05) when only 2 genes are employed. The biological relevance of a set of genes determined by our statistical analysis and the major roles they play in colorectal tumorigenesis is discussed.Conclusions: The method proposed provides statistically significant answers to precise questions relevant for the diagnosis and prognosis of cancer. We found that, with as few as 15 examples, it is possible to train statistically significant classifiers for colon cancer diagnosis. As for the definition of the number of genes sufficient for a reliable classification of colon cancer, our results suggest that it depends on the accuracy required.