Reliable gene signatures for microarray classification:: assessment of stability and performance

Reliable gene signatures for microarray classification:: assessment of stability and performance
复制标题

DOI:
10.1093/bioinformatics/btl400
复制
发表时间:
2006-10-01
期刊:
影响因子:
5.8
通讯作者:
Zimmer, Ralf
Zimmer, Ralf
中科院分区:
生物学3区
文献类型:
--
作者:
Davis, Chad A.;Gerick, Fabian;Zimmer, Ralf

文献摘要

被引文献

相似文献

动机:分析不同样本类别的基因表达测量的两个重要问题是(1)如何对样本进行分类,以及(2)如何识别显示类别和样本子集之间差异的有意义的基因签名(已排序的基因列表)。这两个问题的解决方案都有直接的生物学和生物医学应用。为了获得最佳的分类性能,需要为给定的数据集专门选择合适的分类器和基因选择方法的组合。所选择的基因签名可能不稳定,并且由此产生的分类精度不可靠,特别是在考虑不同的样本子集时。不稳定的基因签名和高估的分类精度都会损害生物学结论。方法:我们通过重复评估所有模型的分类性能来解决这两个问题,即针对随机阵列子集(抽样)的各种基因选择和分类方法的成对组合。模型分数用于为给定数据集选择最合适的模型。一致的基因签名是通过提取那些在多次采样中频繁选择的基因来构建的。抽样还允许测量每个模型的分类性能的稳定性,作为模型可靠性的测量。结果:我们分析了一个包含四个不同软骨样本类别的78个测量值的大型基因表达数据集。在测量子集上训练的分类器经常产生具有高度可变性能的模型。我们的方法通过抽样提供了可靠的分类性能估计。除了可靠的分类性能外,我们还为样本类确定了稳定的共识签名(即基因列表)。人工文献筛选显示,这些基因与我们的骨关节炎软骨基因表达实验高度相关。我们基于可公开获得的乳腺癌数据集,将我们的方法与其他方法进行了比较。
Motivation: Two important questions for the analysis of gene expression measurements from different sample classes are (1) how to classify samples and (2) how to identify meaningful gene signatures (ranked gene lists) exhibiting the differences between classes and sample subsets. Solutions to both questions have immediate biological and biomedical applications. To achieve optimal classification performance, a suitable combination of classifier and gene selection method needs to be specifically selected for a given dataset. The selected gene signatures can be unstable and the resulting classification accuracy unreliable, particularly when considering different subsets of samples. Both unstable gene signatures and overestimated classification accuracy can impair biological conclusions.Methods: We address these two issues by repeatedly evaluating the classification performance of all models, i.e. pairwise combinations of various gene selection and classification methods, for random subsets of arrays (sampling). A model score is used to select the most appropriate model for the given dataset. Consensus gene signatures are constructed by extracting those genes frequently selected over many samplings. Sampling additionally permits measurement of the stability of the classification performance for each model, which serves as a measure of model reliability.Results: We analyzed a large gene expression dataset with 78 measurements of four different cartilage sample classes. Classifiers trained on subsets of measurements frequently produce models with highly variable performance. Our approach provides reliable classification performance estimates via sampling. In addition to reliable classification performance, we determined stable consensus signatures (i.e. gene lists) for sample classes. Manual literature screening showed that these genes are highly relevant to our gene expression experiment with osteoarthritic cartilage. We compared our approach to others based on a publicly available dataset on breast cancer.