WhatsThat? On the Usage of Hierarchical Clustering for Unsupervised Detection & Interpretation of Network Attacks
WhatsThat? On the Usage of Hierarchical Clustering for Unsupervised Detection & Interpretation of Network Attacks
复制标题
DOI:
10.1109/eurospw51379.2020.00084
复制
发表时间:
2020-09
期刊:
影响因子:
--
通讯作者:
Pavol Mulinka;K. Fukuda;P. Casas;L. Kencl
中科院分区:
文献类型:
--
作者:
Pavol Mulinka;K. Fukuda;P. Casas;L. Kencl
The automatic detection and interpretation of network attacks through machine learning is a well-known problem, for which no general solution is available. Super-vised learning and anomaly detection approaches require prior knowledge on the system under analysis, either in the form of normal operation profiles, or on the specific attacks to detect. As a consequence, both approaches have clear limitations when it comes to detecting, and in particular interpreting, previously unseen attacks and anomalies.In this paper we present WhatsThat, a novel approach to unsupervised network anomaly detection, which can both detect and interpret anomalous behaviors in a completely black-box manner, without relying on any ground-truth on the system under analysis. WhatsThat relies on hierarchical–clustering techniques to discover and characterize anomalous patterns present in nested or hierarchically structured multi–dimensional data, which is common in network traffic – e.g., due to multi–layer protocols. The solution uses unsupervised cluster validity metrics to automatically explore the data structure, and builds on automatic identification of relevant features to provide meaningful descriptions for the detected patterns. We showcase WhatsThat in the detection and interpretation of network attacks hidden in real, large-scale network traffic collected at a transit Internet backbone network. While WhatsThat is mainly tailored for unsupervised anomaly detection and interpretation, it can also be applied to the unsupervised analysis of any kind of nested or hierarchically structured multi-dimensional data, showing the potential of hierarchical clustering for general unsupervised data analysis.