A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning

A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning
复制标题

DOI:
10.1109/tkde.2012.35
复制
发表时间:
2013-04-01
影响因子:
8.9
通讯作者:
Herrera, Francisco
Herrera, Francisco
中科院分区:
计算机科学2区
文献类型:
--
作者:
Garcia, Salvador;Luengo, Julian;Herrera, Francisco

文献摘要

被引文献

相似文献

离散化是许多知识发现和数据挖掘任务中必不可少的预处理技术。它的主要目标是将一组连续属性转换为离散属性,通过将分类值与区间相关联,从而将定量数据转换为定性数据。通过这种方式,符号数据挖掘算法可以应用于连续数据,并且简化了信息的表示,使其更加简洁和具体。文献提供了许多离散化的建议,并可以找到一些尝试将它们归类为一个分类法。然而,在以前的论文中,缺乏共识的定义的属性,并没有正式的分类尚未建立,这可能会混淆从业者。此外,只有一小部分离散化方法得到了广泛的考虑,而许多其他方法都没有得到注意。为了缓解这些问题,本文从理论和实证的角度对文献中提出的离散化方法进行了综述。从理论的角度来看,我们开发了一个分类的基础上指出的主要性质在以前的研究中,统一的符号,包括所有已知的方法到目前为止。从经验上讲,我们进行了一个实验研究,在监督分类涉及最具代表性和最新的离散化,不同类型的分类器,和大量的数据集。他们的性能测量的准确性,间隔数和不一致性的结果已通过非参数统计检验进行了验证。此外,一组离散化器被突出显示为性能最好的离散化器。
Discretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones.