Mobile Encrypted Traffic Classification Using Deep Learning

Mobile Encrypted Traffic Classification Using Deep Learning
复制标题

DOI:
10.23919/tma.2018.8506558
复制
发表时间:
2018-06
期刊:
2018 Network Traffic Measurement and Analysis Conference (TMA)
影响因子:
--
通讯作者:
Giuseppe Aceto;D. Ciuonzo;Antonio Montieri;A. Pescapé
Giuseppe Aceto;D. Ciuonzo;Antonio Montieri;A. Pescapé
中科院分区:
其他
文献类型:
--
作者:
Giuseppe Aceto;D. Ciuonzo;Antonio Montieri;A. Pescapé

文献摘要

被引文献

相似文献

手持设备的大规模采用导致了穿越家庭和企业网络以及互联网的移动流量的爆炸性增长。推断产生此类流量的(移动)应用程序的过程称为流量分类(TC),它是高价值配置信息的推动者,但肯定会引发重要的隐私问题。然而,越来越多的加密协议(如TLS)的采用加剧了准确分类器的设计,阻碍了诸如深度包检测等高精度方法的适用性。此外,(每天)不断扩大的应用程序集和移动流量的移动目标性质,使得基于手动和专家发起的功能的常规机器学习设计解决方案过时了。基于这些原因,我们建议将深度学习作为一种可行的策略来设计基于自动提取的特征的流量分类器,以反映复杂的移动流量模式。为此,这里复制、剖析了TC的不同最先进的DL技术,并将其设置为一个系统框架进行比较,其中还包括一个性能评估工作台。基于真实人类用户活动的三个数据集,对这些DL分类器的性能进行了严格的调查,强调了移动加密TC中的DL的陷阱、设计准则和有待解决的问题。
The massive adoption of hand-held devices has led to the explosion of mobile traffic volumes traversing home and enterprise networks, as well as the Internet. Procedures for inferring (mobile) applications generating such traffic, known as Traffic Classification (TC), are the enabler for highly-valuable profiling information while certainly raise important privacy issues. The design of accurate classifiers is however exacerbated by the increasing adoption of encrypted protocols (such as TLS), hindering the applicability of highly-accurate approaches, such as deep packet inspection. Additionally, the (daily) expanding set of apps and the moving-target nature of mobile traffic makes design solutions with usual machine learning, based on manually-and expert-originated features, outdated. For these reasons, we suggest Deep Learning (DL) as a viable strategy to design traffic classifiers based on automatically-extracted features, reflecting the complex mobile-traffic patterns. To this end, different state-of-the-art DL techniques from TC are here reproduced, dissected, and set into a systematic framework for comparison, including also a performance evaluation workbench. Based on three datasets of real human users' activity, performance of these DL classifiers is critically investigated, highlighting pitfalls, design guidelines, and open issues of DL in mobile encrypted TC.