Network Intrusion Detection System Based on an Adversarial Auto-Encoder with Few Labeled Training Samples

Network Intrusion Detection System Based on an Adversarial Auto-Encoder with Few Labeled Training Samples
复制标题

DOI:
10.1007/s10922-022-09698-w
复制
发表时间:
2022-10
影响因子:
3.6
通讯作者:
K. Shiomoto
K. Shiomoto
中科院分区:
计算机科学3区
文献类型:
--
作者:
K. Shiomoto

文献摘要

相似文献

网络入侵检测系统 (NIDS) 对于保护网络系统免受网络攻击至关重要。最近,机器学习已被应用于增强 NIDS 能​​力。为了训练有监督的机器学习模型,需要大量标记的训练样本才能达到实际性能。然而,标记数据样本是一项成本高昂的任务。此外,获取异常数据样本也很困难,因为受到 NIDS 影响的网络流量趋势每天都在变化,并且不断产生新的攻击。为了解决这个问题,我们提出了一种基于半监督机器学习的 NIDS,通过应用对抗性自动编码器(AAE)技术来减少所需的标记训练样本数量。我们通过一系列实验评估了所提出的方法,并证实所提出的基于 AAE 的 NIDS 仅需 0.1% 的标记训练样本即可实现与基于多层感知器的 NIDS 相当的性能。我们还确认,用于注释的数据样本的选择不会影响所提出的基于 AAE 的 NIDS 的性能。我们还评估了所提出方法的性能与其潜变量向量的维数之间的关系。当潜在变量向量的维数为 10 时,通过召回率和 F1 得分衡量的最佳性能出现,这表明该结构可以准确分解攻击和法线。这项研究提出了通过减少标记训练样本数量的半监督学习方法获得的有希望的结果,从而降低了基于机器学习的 NIDS 的运营成本。
Network intrusion detection systems (NIDS) are critical to defending network systems from cyber attacks. Recently, machine learning has been applied to enhance NIDS capability. To train a supervised machine-learning model, a large number of labeled training samples are required to achieve practical performance. However, labeling data samples is a costly task. Additionally, obtaining anomaly data samples is difficult because trends in network traffic that are subject to NIDS change daily, and new attacks continue to be generated. To address this issue, we propose a semi-supervised machine-learning-based NIDS that reduces the required number of labeled training samples by applying an adversarial auto-encoder (AAE) technique. We evaluated the proposed method through a series of experiments and confirmed that the proposed AAE-based NIDS achieves performance comparable to that of multi-layer perceptron-based NIDS with only 0.1% of the labeled training samples. We also confirmed that the selection of data samples for annotation does not affect the performance of the proposed AAE-based NIDS. We also evaluated the relationship between the performance of the proposed method and the dimension of its latent-variable vector. The best performance as measured by recall and F1 score occurred when the dimensionality of the latent variable vector was 10, which suggests that this structure allows for accurate decomposition of attack and normal. This study presents promising results obtained by the proposed semi-supervised learning method with a reduced number of labeled training samples, which reduces the operational costs of a machine-learning-based NIDS.