Deep leakage from gradients

Deep leakage from gradients
复制标题

DOI:
10.48550/arxiv.2301.02621
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Yaqiong Mu
Yaqiong Mu
中科院分区:
其他
文献类型:
--
作者:
Yaqiong Mu

文献摘要

被引文献

相似文献

随着人工智能技术的发展,联邦学习模型以其高效性和保密性在许多行业得到了广泛的应用。一些研究者对其保密性进行了探索,并设计了一些攻击训练数据集的算法,但这些算法都有各自的局限性。因此,大多数人仍然认为本地机器学习梯度信息是安全可靠的。为了引起人们对联邦学习系统安全性的关注,本文设计了一种基于梯度特征的算法对联邦学习模型进行攻击。在联邦学习系统中,梯度与原始训练数据集相比包含的信息很少,但本项目打算通过梯度信息来恢复原始训练图像数据。卷积神经网络(CNN)在图像处理中具有优异的性能。因此,本项目的联邦学习模型采用卷积神经网络结构,并使用图像数据集对模型进行训练。该算法通过生成虚拟图像标签来计算虚拟梯度。然后将虚拟梯度与真实的梯度进行匹配以恢复原始图像。该攻击算法采用Python语言编写,使用猫狗分类Kaggle数据集,并从全连接层逐步扩展到卷积层,从而提高了通用性。目前,该算法恢复的数据与原始图像信息的平均平方误差约为5,绝大多数图像可以根据给定的梯度信息完全恢复,表明联邦学习系统的梯度不是绝对安全可靠的。
With the development of artificial intelligence technology, Federated Learning (FL) model has been widely used in many industries for its high efficiency and confidentiality. Some researchers have explored its confidentiality and designed some algorithms to attack training data sets, but these algorithms all have their own limitations. Therefore, most people still believe that local machine learning gradient information is safe and reliable. In this paper, an algorithm based on gradient features is designed to attack the federated learning model in order to attract more attention to the security of federated learning systems. In federated learning system, gradient contains little information compared with the original training data set, but this project intends to restore the original training image data through gradient information. Convolutional Neural Network (CNN) has excellent performance in image processing. Therefore, the federated learning model of this project is equipped with Convolutional Neural Network structure, and the model is trained by using image data sets. The algorithm calculates the virtual gradient by generating virtual image labels. Then the virtual gradient is matched with the real gradient to restore the original image. This attack algorithm is written in Python language, uses cat and dog classification Kaggle data sets, and gradually extends from the full connection layer to the convolution layer, thus improving the universality. At present, the average squared error between the data recovered by this algorithm and the original image information is approximately 5, and the vast majority of images can be completely restored according to the gradient information given, indicating that the gradient of federated learning system is not absolutely safe and reliable.