Problems of KDD Cup 99 Dataset Existed and Data Preprocessing

Problems of KDD Cup 99 Dataset Existed and Data Preprocessing
复制标题

KDD Cup 99数据集存在的问题及数据预处理

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Huang Jin
Huang Jin
中科院分区:
--
文献类型:
--
作者:
Yan Wang;Kun Yang;Xiang Jing;Huang Jin

文献摘要

被引文献

相似文献

KDD CUP99数据集不仅是入侵检测中应用最广泛的数据集,也是评价入侵检测系统性能优劣的事实上的基准。尽管如此,这个数据集中仍然存在许多不容忽视的问题。为了在入侵检测中建立良好的数据挖掘模型,并找到合适的网络入侵攻击类型的特征,研究人员应该对这些数据集有一个众所周知的了解。本文首先对数据集存在的问题进行了深入的分析,并给出了相关的解决方案。其次,我们还对KDD CUP 99训练集的10%的子集进行了大量的数据预处理,为后续的处理提供了更好的结果。此外,通过对比实验中常见的10种数据挖掘算法,分析总结了数据预处理对数据挖掘算法性能和重要性的重要作用。
KDD Cup 99 dataset is not only the most widely used dataset in intrusion detection, but also the de facto benchmark on evaluating the performance merits of intrusion detection system. Nevertheless there are a lot of issues in this dataset which cannot be omitted. In order to establish good data mining models in intrusion detection and find the appropriate network intrusion attack types’ features, researchers should have a well-known understanding on this dataset. In this paper, first and foremost we have made an in-depth analysis on the problems which the dataset are existed, and given the related solutions. Secondly, we also have carried out plenty data preprocessing on the 10% subset of KDD Cup 99 dataset’s training set, giving better results to the following process. What’s more, by comparing 10 common kinds of data mining algorithms in our experiment, we have analyzed and summarized that data preprocessing plays a vital role on the performance and importance to data mining algorithms.