FedPacket: A Federated Learning Approach to Mobile Packet Classification

FedPacket: A Federated Learning Approach to Mobile Packet Classification
复制标题

DOI:
10.1109/tmc.2021.3058627
复制
发表时间:
2022-10-01
影响因子:
7.9
通讯作者:
Markopoulou, Athina
Markopoulou, Athina
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bakopoulou, Evita;Tillman, Balint;Markopoulou, Athina

文献摘要

被引文献

相似文献

为了提高移动数据的透明度,人们提出了各种方法来检查移动设备产生的网络流量,并检测个人身份信息(PII)、广告请求等的暴露。最先进的方法使用从HTTP数据包中提取的特征,并以集中的方式训练分类器:用户在其移动设备上收集和标记网络数据包,然后将数据上传到中央服务器;服务器使用所有用户提供的数据来训练数据包分类器。但是,从用户设备上收集的网络流量的训练数据集可能包含用户可能不想上传的敏感信息。在本文中,我们提出了一种用于移动数据包分类的联邦学习方法,该方法使设备能够协作训练全局模型,而无需将收集到的训练数据上传到设备上。我们将我们的框架应用于两个数据包分类任务(即,预测单个数据包中的PII暴露或广告请求),并使用三个真实世界的数据集,在分类性能,通信和计算成本方面证明了其有效性。我们在这个过程中解决的方法挑战包括模型和特征选择,以及为我们的包分类任务专门调优联邦学习参数。我们还讨论了隐私限制和缓解方法。
In order to improve mobile data transparency, various approaches have been proposed to inspect network traffic generated by mobile devices and detect exposure of personally identifiable information (PII), ad requests, etc. State-of-the-art approaches use features extracted from HTTP packets and train classifiers in a centralized way: users collect and label network packets on their mobile devices, then upload data to a central server; the server uses the data contributed by all users to train a packet classifier. However, training datasets from network traffic collected on user devices may contain sensitive information that users may not want to upload. In this article, we propose a federated learning approach to mobile packet classification, which enables devices to collaboratively train a global model, without uploading the training data collected on devices. We apply our framework to two packet classification tasks (i.e., to predict PII exposure or ad requests in individual packets) and we demonstrate its effectiveness in terms of classification performance, communication and computation cost, using three real-world datasets. Methodological challenges we address in the process include model and feature selection, as well as tuning the federated learning parameters specifically for our packet classification tasks. We also discuss privacy limitations and mitigation approaches.