On the Performance of Machine Learning Models for Anomaly-Based Intelligent Intrusion Detection Systems for the Internet of Things

On the Performance of Machine Learning Models for Anomaly-Based Intelligent Intrusion Detection Systems for the Internet of Things
复制标题

DOI:
10.1109/jiot.2021.3103829
复制
发表时间:
2022-03-15
影响因子:
10.6
通讯作者:
Rahman, Abdul
Rahman, Abdul
中科院分区:
计算机科学1区
文献类型:
--
作者:
Abdelmoumin, Ghada;Rawat, Danda B.;Rahman, Abdul

文献摘要

被引文献

相似文献

与基于深度学习的入侵检测系统(DL-IDS)相比,基于异常的机器学习入侵检测系统(AML-IDS)在检测物联网(IoT)中的入侵时表现出较低的性能和预测准确性。特别地,与具有两类神经网络(2-NN)方法的DL-IDS相比,采用用于IoT的低复杂度模型(诸如主成分机(PCA)方法和单类支持向量机(1-SVM)方法)的AML-IDS在检测入侵方面是低效的。与DL-IDS相比,PCA和1-SVM AML-IDS遭受低检测率。与DL-IDS相比,数据集的大小和数据集中的特征或变体的数量可能会影响PCA和1-SVM AML-IDS的性能。我们把AML-IDS模型的低性能和预测准确性归因于不平衡的数据集、训练数据和测试数据之间的低相似性指数以及使用单学习者模型。单学习者模型的固有局限性直接影响了智能入侵检测系统的准确性。此外,测试数据和训练数据之间的不相似性导致AML-IDS中的假阳性(FP)率比DL-IDS越来越高,DL-IDS具有低误报和高可预测性。在本文中,我们研究了使用优化技术来提高单学习者AML-IDS的性能,例如PCA和1-SVM AML-IDS模型,用于构建高效,可扩展和分布式的智能IDS,以检测物联网中的入侵。我们通过使用Microsoft Azure ML Studio(AMLS)平台和包含恶意和良性物联网和工业物联网(IIoT)网络流量的两个数据集调整超参数和集成学习优化技术来评估这些AML-IDS模型。此外,我们还对物联网AML-IDS模型的性能和可预测性进行了比较分析。
Anomaly-based machine learning-enabled intrusion detection systems (AML-IDSs) show low performance and prediction accuracy while detecting intrusions in the Internet of Things (IoT) than that of deep learning-based intrusion detection systems (DL-IDSs). In particular, AML-IDS that employ low complexity models for IoT, such as the principal component machine (PCA) method and the one-class support vector machine (1-SVM) method, are inefficient in detecting intrusions when compared to DL-IDS with the two-class neural network (2-NN) method. PCA and 1-SVM AML-IDS suffer from low detection rates compared to DL-IDS. The size of the data set and the number of features or variants in the data set may influence how well PCA and 1-SVM AML-IDS perform compared to DL-IDS. We attribute the low performance and prediction accuracy of the AML-IDS model to an imbalanced data set, a low similarity index between the training data and testing data, and the use of a single-learner model. The intrinsic limitations of the single-learner model have a direct impact on the accuracy of an intelligent IDS. Also, the dissimilarity between testing data and training data leads to an increasingly high rate of false positives (FPs) in AML-IDS than DL-IDS, which have low false alarms and high predictability. In this article, we examine the use of optimization techniques to enhance the performance of single-learner AML-IDS, such as PCA and 1-SVM AML-IDS models for building efficient, scalable, and distributed intelligent IDS for detecting intrusions in IoT. We evaluate these AML-IDS models by tuning hyperparameters and ensemble learning optimization techniques using the Microsoft Azure ML Studio (AMLS) platform and two data sets containing malicious and benign IoT and industrial IoT (IIoT) network traffic. Furthermore, we present a comparative analysis of AML-IDS models for IoT regarding their performance and predictability.