DAHID: Domain Adaptive Host-based Intrusion Detection

DAHID: Domain Adaptive Host-based Intrusion Detection
复制标题

DOI:
10.1109/csr51186.2021.9527966
复制
发表时间:
2021-07
期刊:
2021 IEEE International Conference on Cyber Security and Resilience (CSR)
影响因子:
--
通讯作者:
Oluwagbemiga Ajayi;A. Gangopadhyay
Oluwagbemiga Ajayi;A. Gangopadhyay
中科院分区:
其他
文献类型:
--
作者:
Oluwagbemiga Ajayi;A. Gangopadhyay

文献摘要

相似文献

随着越来越多的网络物理系统被部署,网络安全变得越来越重要,攻击面也随之爆炸。由于收集标记数据和训练模型的成本,为每个计算基础设施和各种攻击场景创建具有可接受性能的模型是不切实际的。因此,重要的是要能够开发模型,可以利用现有的知识在攻击源域,以提高性能在目标域与域的具体data.In这项工作中,我们提出了域自适应基于主机的入侵检测DAHID;一种方法,用于检测攻击在多个域的网络安全。具体来说,我们实现了一个深度学习模型,该模型利用了少量的目标域数据进行基于主机的入侵检测。在我们的实验中,我们使用了来自澳大利亚国防军学院的两个数据集; ADFA-WD作为源域,ADFA-WD:SAA作为目标域数据集。当我们用20%的ADFA-WD:SAA微调在ADFA-WD上训练的深度学习模型时,我们记录到曲线下面积AUC从83%显著改善到91%。我们的研究结果表明,迁移学习可以帮助减轻在构建基于主机的入侵检测模型中对大量特定领域数据集的需求。
Cybersecurity is becoming increasingly important with the explosion of attack surfaces as more cyber-physical systems are being deployed. It is impractical to create models with acceptable performance for every single computing infrastructure and the various attack scenarios due to the cost of collecting labeled data and training models. Hence it is important to be able to develop models that can take advantage of knowledge available in an attack source domain to improve performance in a target domain with little domain specific data.In this work we proposed Domain Adaptive Host-based Intrusion Detection DAHID; an approach for detecting attacks in multiple domains for cybersecurity. Specifically, we implemented a deep learning model which utilizes a substantially smaller amount of target domain data for host-based intrusion detection.In our experiments, we used two datasets from Australian Defense Force Academy; ADFA-WD as the source domain and ADFA-WD:SAA as the target domain datasets. We recorded a significant improvement in Area Under Curve AUC from 83% to 91%, when we fine-tuned a deep learning model trained on ADFA-WD with as little as 20% of ADFA-WD:SAA. Our result shows transfer learning can help to alleviate the need of huge domain specific dataset in building host-based intrusion detection models.