Comparative analysis of instance selection algorithms for instance-based classifiers in the context of medical decision support

Comparative analysis of instance selection algorithms for instance-based classifiers in the context of medical decision support
复制标题

DOI:
10.1088/0031-9155/56/2/012
复制
发表时间:
2011-01-21
影响因子:
3.5
通讯作者:
Tourassi, Georgia D.
Tourassi, Georgia D.
中科院分区:
工程技术2区
文献类型:
--
作者:
Mazurowski, Maciej A.;Malof, Jordan M.;Tourassi, Georgia D.

文献摘要

被引文献

相似文献

在构建模式分类器时,充分利用可用于其开发的实例(也称为案例、示例、模式或原型)非常重要。在本文中,我们对算法进行了广泛的比较分析,在给定先前获取的实例池的情况下,尝试选择在分类性能、时间效率和存储要求方面最有效地构建基于实例的分类器的算法。我们评估了之前提出的七种实例选择算法,并将它们的性能与简单的实例随机选择进行了比较。我们使用 k-近邻分类器和三个分类问题进行评估:一个使用模拟高斯数据,两个分别基于乳腺癌检测和诊断的临床数据库。最后,我们评估可供选择的实例数量对选择算法性能的影响,并对所选实例进行初步分析。实验表明,对于所有研究的分类问题,可以将原始开发数据集的大小减少到其初始大小的 3% 以下,同时保持或提高分类性能。随机突变爬山成为更好的选择算法。此外,我们表明一些先前提出的算法的性能比随机选择更差。关于可用于分类器开发的实例数量对选择算法性能的影响,我们确认,随着可用实例池的增加,选择算法通常更有效。总之,实例选择通常对于基于实例的分类器是有益的,因为它可以提高其性能,减少其存储需求并缩短其响应时间。然而,选择正确的选择算法至关重要。
When constructing a pattern classifier, it is important to make best use of the instances (a.k.a. cases, examples, patterns or prototypes) available for its development. In this paper we present an extensive comparative analysis of algorithms that, given a pool of previously acquired instances, attempt to select those that will be the most effective to construct an instance-based classifier in terms of classification performance, time efficiency and storage requirements. We evaluate seven previously proposed instance selection algorithms and compare their performance to simple random selection of instances. We perform the evaluation using k-nearest neighbor classifier and three classification problems: one with simulated Gaussian data and two based on clinical databases for breast cancer detection and diagnosis, respectively. Finally, we evaluate the impact of the number of instances available for selection on the performance of the selection algorithms and conduct initial analysis of the selected instances. The experiments show that for all investigated classification problems, it was possible to reduce the size of the original development dataset to less than 3% of its initial size while maintaining or improving the classification performance. Random mutation hill climbing emerges as the superior selection algorithm. Furthermore, we show that some previously proposed algorithms perform worse than random selection. Regarding the impact of the number of instances available for the classifier development on the performance of the selection algorithms, we confirm that the selection algorithms are generally more effective as the pool of available instances increases. In conclusion, instance selection is generally beneficial for instance-based classifiers as it can improve their performance, reduce their storage requirements and improve their response time. However, choosing the right selection algorithm is crucial.