Identifying Encrypted Malware Traffic with Contextual Flow Data

Identifying Encrypted Malware Traffic with Contextual Flow Data
复制标题

DOI:
10.1145/2996758.2996768
复制
发表时间:
2016-10
期刊:
Proceedings of the 2016 ACM Workshop on Artificial Intelligence and Security
影响因子:
--
通讯作者:
Blake Anderson;D. McGrew
Blake Anderson;D. McGrew
中科院分区:
其他
文献类型:
--
作者:
Blake Anderson;D. McGrew

文献摘要

被引文献

相似文献

识别加密网络流量中包含的威胁是一组独特的挑战。重要的是要监控此流量的威胁和恶意软件,但这样做的方式,以保持加密的完整性。因为模式匹配不能对加密数据进行操作,所以先前的方法已经利用了从流收集的可观察元数据,例如,流的分组长度和到达间隔时间。在这项工作中,我们扩展了目前的国家的最先进的考虑数据Omnia的方法。为此,我们开发了有监督的机器学习模型,这些模型利用了一组独特而多样的网络流数据特征。这些数据特征包括TLS握手元数据、链接到加密流的DNS上下文流以及在5分钟窗口内来自相同源IP地址的HTTP上下文流的HTTP报头。我们开始展示恶意和良性流量之间的差异使用TLS,DNS和HTTP上数以百万计的独特的流。这项研究是用来设计的功能集,具有最大的歧视的权力。然后,我们表明,将此上下文信息纳入监督学习系统显着提高性能在0.00%的错误发现率的问题进行分类加密,恶意流量。我们进一步验证了我们的假阳性率在一个独立的,真实世界的数据集。
Identifying threats contained within encrypted network traffic poses a unique set of challenges. It is important to monitor this traffic for threats and malware, but do so in a way that maintains the integrity of the encryption. Because pattern matching cannot operate on encrypted data, previous approaches have leveraged observable metadata gathered from the flow, e.g., the flow's packet lengths and inter-arrival times. In this work, we extend the current state-of-the-art by considering a data omnia approach. To this end, we develop supervised machine learning models that take advantage of a unique and diverse set of network flow data features. These data features include TLS handshake metadata, DNS contextual flows linked to the encrypted flow, and the HTTP headers of HTTP contextual flows from the same source IP address within a 5 minute window. We begin by exhibiting the differences between malicious and benign traffic's use of TLS, DNS, and HTTP on millions of unique flows. This study is used to design the feature sets that have the most discriminatory power. We then show that incorporating this contextual information into a supervised learning system significantly increases performance at a 0.00% false discovery rate for the problem of classifying encrypted, malicious flows. We further validate our false positive rate on an independent, real-world dataset.