Genetic Programming for Automatically Constructing Data Mining Algorithms

Genetic Programming for Automatically Constructing Data Mining Algorithms
复制标题

自动构建数据挖掘算法的遗传编程

DOI:
--
复制
发表时间:
2009
期刊:
Encyclopedia of Data Warehousing and Mining
影响因子:
--
通讯作者:
G. Pappa
G. Pappa
中科院分区:
--
文献类型:
--
作者:
A. Freitas;G. Pappa

文献摘要

被引文献

相似文献

目前,研究人员和实践者可以使用多种数据挖掘算法(Witten & Frank,2005;Tan 等人,2006)。尽管这些算法千差万别,但几乎所有算法都有一个共同特征:它们都是手动设计的。因此,当前的数据挖掘算法通常在其设计中融入了人类的偏见和先入之见。本文提出了数据挖掘算法设计的另一种方法,即通过遗传编程(GP)自动创建数据挖掘算法(Pappa & Freitas,2006)。从本质上讲,GP 是一种进化算法,即一种受达尔文自然选择过程启发的搜索算法,它会进化计算机程序或可执行结构。这种方法开辟了新的研究途径,提供了设计新颖的数据挖掘算法的方法,这些算法较少受到人类偏见和先入之见的限制,因此为用户提供了发现新类型模式(或知识)的潜力。它还为自动创建针对正在挖掘的数据量身定制的数据挖掘算法提供了一个有趣的机会。
At present there is a wide range of data mining algorithms available to researchers and practitioners (Witten & Frank, 2005; Tan et al., 2006). Despite the great diversity of these algorithms, virtually all of them share one feature: they have been manually designed. As a result, current data mining algorithms in general incorporate human biases and preconceptions in their designs. This article proposes an alternative approach to the design of data mining algorithms, namely the automatic creation of data mining algorithms by means of Genetic Programming (GP) (Pappa & Freitas, 2006). In essence, GP is a type of Evolutionary Algorithm – i.e., a search algorithm inspired by the Darwinian process of natural selection – that evolves computer programs or executable structures. This approach opens new avenues for research, providing the means to design novel data mining algorithms that are less limited by human biases and preconceptions, and so offer the potential to discover new kinds of patterns (or knowledge) to the user. It also offers an interesting opportunity for the automatic creation of data mining algorithms tailored to the data being mined.