A New Implementation of Federated Learning for Privacy and Security Enhancement

A New Implementation of Federated Learning for Privacy and Security Enhancement
复制标题

DOI:
10.1109/globecom48099.2022.10001614
复制
发表时间:
2022-08
期刊:
GLOBECOM 2022 - 2022 IEEE Global Communications Conference
影响因子:
--
通讯作者:
Xiang Ma;Haijian Sun;R. Hu;Yi Qian
Xiang Ma;Haijian Sun;R. Hu;Yi Qian
中科院分区:
其他
文献类型:
--
作者:
Xiang Ma;Haijian Sun;R. Hu;Yi Qian

文献摘要

相似文献

由于对个人数据隐私的日益关注以及本地客户端数据量的快速增长,联邦学习(FL)已经成为一种新的机器学习环境。FL系统由一个中央参数服务器和多个本地客户端组成。它将数据保存在本地客户端,并通过共享本地学习的模型参数来学习集中式模型。无需共享本地数据,隐私可以得到很好的保护。然而,由于共享的是模型而不是原始数据,因此系统可能会受到恶意客户端发起的中毒模型攻击。此外,由于服务器上没有可用的本地客户端数据,因此识别恶意客户端具有挑战性。此外,成员推理攻击仍然可以通过使用上传的模型来估计客户端的本地数据,导致隐私泄露。在这项工作中,我们首先提出了一个基于模型更新的联邦平均算法,以抵御拜占庭攻击,如加性噪声攻击和符号翻转攻击。提出了个体客户端模型初始化方法,通过隐藏个体本地机器学习模型来提供进一步的隐私保护以免受成员推断攻击。当结合这两种方案时,可以有效地增强隐私和安全性。实验证明,所提出的方案收敛于非IID数据分布时,没有攻击。在Byzantine攻击下,该方案的性能优于经典的基于模型的FedAvg算法.
Motivated by the ever-increasing concerns on per-sonal data privacy and the rapidly growing data volume at local clients, federated learning (FL) has emerged as a new machine learning setting. An FL system is comprised of a central parame-ter server and multiple local clients. It keeps data at local clients and learns a centralized model by sharing the model parameters learned locally. No local data needs to be shared, and privacy can be well protected. Nevertheless, since it is the model instead of the raw data that is shared, the system can be exposed to the poisoning model attacks launched by malicious clients. Furthermore, it is challenging to identify malicious clients since no local client data is available on the server. Besides, membership inference attacks can still be performed by using the uploaded model to estimate the client's local data, leading to privacy disclosure. In this work, we first propose a model update based federated averaging algorithm to defend against Byzantine attacks such as additive noise attacks and sign-flipping attacks. The individual client model initialization method is presented to provide further privacy protections from the membership inference attacks by hiding the individual local machine learning model. When combining these two schemes, privacy and security can be both effectively enhanced. The proposed schemes are proved to converge experimentally under non-IID data distribution when there are no attacks. Under Byzantine attacks, the proposed schemes perform much better than the classical model based FedAvg algorithm.