Win percentage: a novel measure for assessing the suitability of machine classifiers for biological problems.

Win percentage: a novel measure for assessing the suitability of machine classifiers for biological problems.
复制标题

DOI:
10.1186/1471-2105-13-s3-s7
复制
发表时间:
2012-03-21
期刊:
影响因子:
3
通讯作者:
Wang MD
Wang MD
中科院分区:
生物学4区
文献类型:
--
作者:
Parry RM;Phan JH;Wang MD

文献摘要

相似文献

为特定的生物学应用选择合适的分类器对研究人员和从业者都提出了难题。特别地,选择分类器在很大程度上取决于所选择的特征。对于高吞吐量的生物医学数据集,特征选择通常是一个预处理步骤,它为使用相同建模假设构建的分类器提供了不公平的优势。在本文中,我们寻求适合于一个特定的问题独立的特征选择的分类器。我们提出了一种新的措施,称为“赢的百分比”,用于评估机器分类器的适用性,以一个特定的问题。我们将获胜百分比定义为分类器在有限的随机特征集样本上表现得比同行更好的概率,使每个分类器都有平等的机会找到合适的特征。首先,我们说明了在特征选择之后评估分类器的困难。我们发现,几个分类器可以在统计上显着优于他们的同行给予正确的功能集之间的前0.001%的所有功能集。我们说明了使用合成数据的获胜率的效用,并在分析代表三种疾病的八个微阵列数据集:乳腺癌,多发性骨髓瘤和神经母细胞瘤中评估六个分类器。在最初使用所有高斯基因对之后,我们表明可以使用所有特征对的较小随机样本来实现获胜百分比的精确估计(在1%以内)。我们表明,对于这些数据,没有一个分类器可以被认为是最好的,而不知道的特征集。相反,获胜百分比捕获了每个分类器将基于对性能的经验估计优于其同行的非零概率。从根本上说,我们说明了最合适的分类器的选择(即,一个更有可能比它的对等体表现得更好)不仅取决于数据集和应用程序,而且还取决于特征选择的彻底性。特别是,获胜百分比提供了一个单一的测量,可以帮助用户消除或选择分类器,为他们的特定应用程序。
Selecting an appropriate classifier for a particular biological application poses a difficult problem for researchers and practitioners alike. In particular, choosing a classifier depends heavily on the features selected. For high-throughput biomedical datasets, feature selection is often a preprocessing step that gives an unfair advantage to the classifiers built with the same modeling assumptions. In this paper, we seek classifiers that are suitable to a particular problem independent of feature selection. We propose a novel measure, called "win percentage", for assessing the suitability of machine classifiers to a particular problem. We define win percentage as the probability a classifier will perform better than its peers on a finite random sample of feature sets, giving each classifier equal opportunity to find suitable features. First, we illustrate the difficulty in evaluating classifiers after feature selection. We show that several classifiers can each perform statistically significantly better than their peers given the right feature set among the top 0.001% of all feature sets. We illustrate the utility of win percentage using synthetic data, and evaluate six classifiers in analyzing eight microarray datasets representing three diseases: breast cancer, multiple myeloma, and neuroblastoma. After initially using all Gaussian gene-pairs, we show that precise estimates of win percentage (within 1%) can be achieved using a smaller random sample of all feature pairs. We show that for these data no single classifier can be considered the best without knowing the feature set. Instead, win percentage captures the non-zero probability that each classifier will outperform its peers based on an empirical estimate of performance. Fundamentally, we illustrate that the selection of the most suitable classifier (i.e., one that is more likely to perform better than its peers) not only depends on the dataset and application but also on the thoroughness of feature selection. In particular, win percentage provides a single measurement that could assist users in eliminating or selecting classifiers for their particular application.