Sampling Methods in Genetic Programming Learners from Large Datasets: A Comparative Study

Sampling Methods in Genetic Programming Learners from Large Datasets: A Comparative Study
复制标题

大数据集中遗传编程学习者的抽样方法:比较研究

DOI:
--
复制
发表时间:
2016
期刊:
INNS Conference on Big Data
影响因子:
--
通讯作者:
M. Rukoz
M. Rukoz
中科院分区:
--
文献类型:
--
作者:
Hmida Hmida;S. B. Hamida;A. Borgi;M. Rukoz

文献摘要

被引文献

相似文献

随着大数据时代的到来,可用于数据挖掘和知识发现的数据量继续快速增长。遗传编程算法作为一种高效的机器学习技术,面临着处理海量数据的新挑战。主动采样已经用于主动学习,它可能是从非常大的数据集改进进化算法(EA)训练的一个很好的解决方案。本文回顾了主动GP学习者已经使用的抽样技术,并讨论了它们从超大数据集改进GP训练的能力。在每个采样策略中实现了一种方法,并将其应用于使用非常接近的参数的KDD入侵检测问题。实验结果表明,采样方法的性能优于完全数据集的结果,但其中一些方法不能扩展到大数据集。
The amount of available data for data mining and knowledge discovery continue to grow very fast with the era of Big Data. Genetic Programming algorithms (GP), that are efficient machine learning techniques, are face up to a new challenge that is to deal with the mass of the provided data. Active Sampling, already used for Active Learning, might be a good solution to improve the Evolutionary Algorithms (EA) training from very big data sets. This paper present a review of sampling techniques already used with active GP learner and discuss their ability to improve the GP training from very big data sets. A method in each sampling strategy is implemented and applied on the KDD intrusion detection problem using very close parameters. Experimental results show that sampling methods outperforms results obtained with full dataset but some of them cannot be scaled to large datasets.