Recompose Event Sequences vs. Predict Next Events: A Novel Anomaly Detection Approach for Discrete Event Logs

Recompose Event Sequences vs. Predict Next Events: A Novel Anomaly Detection Approach for Discrete Event Logs
复制标题

DOI:
10.1145/3433210.3453098
复制
发表时间:
2020-12
期刊:
Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Lun-Pin Yuan;Peng Liu;Sencun Zhu
Lun-Pin Yuan;Peng Liu;Sencun Zhu
中科院分区:
其他
文献类型:
--
作者:
Lun-Pin Yuan;Peng Liu;Sencun Zhu

文献摘要

相似文献

离散事件日志的异常检测是入侵检测领域最具挑战性的问题之一。虽然大多数早期的工作集中在将无监督学习应用于工程特征,但最近的工作已经开始通过将深度学习方法应用于离散事件条目的抽象来解决这一挑战。受自然语言处理的启发,提出了基于LSTM的异常检测模型。它们试图预测即将发生的事件,并在预测不符合特定标准时发出异常警报。然而,这样的预测下一个事件的方法有一个基本的限制:事件预测可能无法充分利用序列的独特特征。这种限制导致高假阳性(FP)。同样重要的是要检查序列的结构和单个事件之间的双向因果关系。为此,我们提出了一种新的方法:重组事件序列作为异常检测。我们提出了DabLog,一种基于LSTM的基于深度自动编码器的离散事件异常检测方法。根本区别在于,我们的方法不是预测即将发生的事件,而是通过分析(编码)和重建(解码)给定序列来确定序列是正常还是异常。我们的评估结果表明,我们的新方法可以显着减少FP的数量,从而实现更高的F1分数。
One of the most challenging problems in the field of intrusion detection is anomaly detection for discrete event logs. While most earlier work focused on applying unsupervised learning upon engineered features, most recent work has started to resolve this challenge by applying deep learning methodology to abstraction of discrete event entries. Inspired by natural language processing, LSTM-based anomaly detection models were proposed. They try to predict upcoming events, and raise an anomaly alert when a prediction fails to meet a certain criterion. However, such a predict-next-event methodology has a fundamental limitation: event predictions may not be able to fully exploit the distinctive characteristics of sequences. This limitation leads to high false positives (FPs). It is also critical to examine the structure of sequences and the bi-directional causality among individual events. To this end, we propose a new methodology: Recomposing event sequences as anomaly detection. We propose DabLog, a LSTM-based Deep Autoencoder-Based anomaly detection method for discrete event Logs. The fundamental difference is that, rather than predicting upcoming events, our approach determines whether a sequence is normal or abnormal by analyzing (encoding) and reconstructing (decoding) the given sequence. Our evaluation results show that our new methodology can significantly reduce the numbers of FPs, hence achieving a higher F1 score.