Mobile app identification for encrypted network flows by traffic correlation

Mobile app identification for encrypted network flows by traffic correlation
复制标题

DOI:
10.1177/1550147718817292
复制
发表时间:
2018-12
影响因子:
2.3
通讯作者:
Gaofeng He;Bingfeng Xu;Lu Zhang;Haiting Zhu
Gaofeng He;Bingfeng Xu;Lu Zhang;Haiting Zhu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Gaofeng He;Bingfeng Xu;Lu Zhang;Haiting Zhu

文献摘要

相似文献

每流粒度的移动应用程序(简称“app”)识别对于流量工程、网络管理和安全实践至关重要。然而,越来越多的加密流量(如超文本传输协议安全)会造成不确定性。为了应对这一挑战,我们仔细分析了移动应用流量(主要包括域名系统、超文本传输协议和加密流量,如安全套接字层和传输层安全),并观察到(1)不同应用查询的服务器主机名集是可区分的;(2)移动应用可能同时查询多个服务器主机名,即应用可能在短时间间隔内发送多个域名系统查找;以及(3)加密流量可能类似于同一应用产生的其他各种网络流量。基于这三点,本文提出了一种新的针对加密网络流量的APP识别方法。具体地说,研究了时间、词汇和元数据的相似性来选择相关的流量,并采用信息检索技术来识别应用程序。我们运行了一组全面的实验来评估所提出的方法的性能。实验结果表明,该方法识别正确率可高达95%,且具有存储需求小、训练速度快等优点。
Mobile application (simply “app”) identification at a per-flow granularity is vital for traffic engineering, network management, and security practices. However, uncertainty is caused by a growing fraction of encrypted traffic such as Hypertext Transfer Protocol Secure. To address this challenge, we have carefully analyzed mobile app traffic (mainly including Domain Name System, Hypertext Transfer Protocol, and encrypted traffic such as Secure Sockets Layer and Transport Layer Security) and observed that (1) the sets of server hostnames queried by different apps are distinguishable; (2) mobile apps may query multiple server hostnames simultaneously, that is, apps may send several Domain Name System lookups within a short time interval; and (3) the encrypted traffic may be similar to various other network flows generated by the same app. Based on these three observations, in this article, we propose a novel app identification methodology for encrypted network flows. To be specific, temporal, lexical, and metadata similarity are investigated to select correlated traffic and information retrieving techniques are adopted to identify apps. We ran a thorough set of experiments to assess the performance of the proposed approaches. The experimental results show that the identification accuracy can be as high as 95%, and the proposed methods have low storage requirements as well as fast training speeds.