Efficacy of Machine Learning-Based Classifiers for Binary and Multi-Class Network Intrusion Detection

Efficacy of Machine Learning-Based Classifiers for Binary and Multi-Class Network Intrusion Detection
复制标题

基于机器学习的分类器在二元和多类网络入侵检测中的功效

DOI:
10.1109/i2cacis52118.2021.9495877
复制
发表时间:
2021
期刊:
2021 IEEE International Conference on Automatic Control & Intelligent Systems (I2CACIS)
影响因子:
--
通讯作者:
M. Chouikha
M. Chouikha
中科院分区:
--
文献类型:
--
作者:
Toya Acharya;Ishan Khatri;A. Annamalai;M. Chouikha

文献摘要

被引文献

相似文献

基于互联网的服务无疑以指数级的增长引领了全球革命,但安全漏洞导致个人数字资产损失,需要全面的网络安全解决方案。传统的基于特征的网络入侵检测方法是用来捕获网络中正常和异常流量的属性,但它不能检测零日攻击。在众多已知的网络入侵检测方法中,基于机器学习的方法能够有效地分析大量的网络流量数据,有效地检测零日攻击,因此受到了广泛的关注。在实际实施方案中,不平衡的NIDS数据集不能提供更好的性能。将目标类的数量减少为新的目标类创建了平衡的NID和改进的分类器性能。在本文中,我们介绍了几种机器学习算法的有效性,包括随机森林(RF),J48,朴素贝叶斯,贝叶斯网络,袋,AdaBoost,和支持向量机(支持向量机)使用网络日志流量(KDD99,UNSW-NB15,和CIC-IDS2017)使用WEKA。本文研究了改变公开可用的网络入侵数据集的输出类别数对敏感度(真阳性率)、假阳性率(FPR)、ROC曲线下面积(AUC)和错误识别率的影响。有趣的是,这些分类器的效率已经提高,为目标类添加了强关联的特征。实验结果表明,机器学习分类器的性能随着目标类个数的减少而提高。向输出类添加高度相关的特征提高了分类器的性能。
The internet-based services undoubtedly led the worldwide revolution with exponential growth, but security breaches resulting personal digital asset losses which need for a comprehensive cybersecurity solution. Traditionally, signature-based network intrusion detection is employed to capture attributes of normal and abnormal traffics in a network, but it fails to detect the zero-day attack. The machine learning-based approach is attractive among various known NIDS methods to circumvent the shortcoming because machine learning based approach can efficiently analyze the big network traffic data and efficiently detect the zero-day attack. The imbalanced NIDS dataset does not provide better performance on practical implementation scenarios. Reducing the number of target classes into a new target class creates a balanced NIDS and improved classifier performance. In this paper, we present the efficacy of several machine learning algorithms, including Random forest (RF), J48, Naïve Bayes, Bayesian Network, Bagging, AdaBoost, and Support Vector Machine (SVM) using network logs traffic (KDD99, UNSW-NB15, and CIC-IDS2017) using WEKA. This paper examined the impact of changing the number of output classes of the publicly available network intrusion datasets on sensitivity (True Positive Rate), False Positive Rate (FPR), Area under the ROC curve (AUC) and incorrectly identified percentage. Interestingly, the efficiency of these classifiers has increased, adding strongly correlated features to the target classes. The experimented results reveal that the machine learning classifiers performance improved when the number of target classes decreased. The addition of a highly correlated feature to the output class increases the performance of the classifiers.