Automatic recommendation of classification algorithms based on data set characteristics

Automatic recommendation of classification algorithms based on data set characteristics
复制标题

基于数据集特征的分类算法自动推荐

DOI:
10.1016/j.patcog.2011.12.025
复制
发表时间:
2012-07-01
影响因子:
8
通讯作者:
Wang, Chao
Wang, Chao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Song, Qinbao;Wang, Guangtao;Wang, Chao

文献摘要

被引文献

相似文献

为给定的数据集选择合适的分类算法在实践中是非常重要和有用的,但也充满了挑战。本文提出了一种分类算法推荐方法。首先,使用一种新的方法提取数据集的特征向量,并对分类算法在数据集上的性能进行评估。然后提取新数据集的特征向量,并识别其k个最近的数据集。然后,将最近的数据集的分类算法推荐给新的数据集。提出的数据集特征提取方法使用结构和统计信息来表征数据集,这与现有的方法有很大的不同。为了评估所提出的分类算法推荐方法和数据集特征提取方法的性能,在84个公开可用的UCI数据集上对17种不同类型的分类算法、3种不同类型的数据集特征化方法和所有可能数量的最近数据集进行了广泛的实验。结果表明,该方法是有效的,可以在实际中使用。(C)2012爱思唯尔有限公司保留所有权利。
Choosing appropriate classification algorithms for a given data set is very important and useful in practice but also is full of challenges. In this paper, a method of recommending classification algorithms is proposed. Firstly the feature vectors of data sets are extracted using a novel method and the performance of classification algorithms on the data sets is evaluated. Then the feature vector of a new data set is extracted, and its k nearest data sets are identified. Afterwards, the classification algorithms of the nearest data sets are recommended to the new data set. The proposed data set feature extraction method uses structural and statistical information to characterize data sets, which is quite different from the existing methods. To evaluate the performance of the proposed classification algorithm recommendation method and the data set feature extraction method, extensive experiments with the 17 different types of classification algorithms, the three different types of data set characterization methods and all possible numbers of the nearest data sets are conducted upon the 84 publicly available UCI data sets. The results indicate that the proposed method is effective and can be used in practice. (C) 2012 Elsevier Ltd. All rights reserved.