Distributed generation of privacy preserving data with user customization

Distributed generation of privacy preserving data with user customization
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiao Chen;Thomas Navidi;Stefano Ermon;R. Rajagopal
Xiao Chen;Thomas Navidi;Stefano Ermon;R. Rajagopal
中科院分区:
其他
文献类型:
--
作者:
Xiao Chen;Thomas Navidi;Stefano Ermon;R. Rajagopal

文献摘要

相似文献

分布式设备(如移动的电话)可以产生和存储大量数据,这些数据可以增强机器学习模型;但是,这些数据可能包含特定于数据所有者的私有信息,从而阻止数据的发布。我们希望在保持有用信息的同时,减少用户特定的私人信息和数据之间的相关性。而不是学习一个大的模型,以实现从端到端的私有化,我们引入了一个解耦的创建一个潜在的表示和私有化的数据,允许用户特定的私有化发生在一个分布式设置有限的计算和最小的干扰效用的数据。我们利用变分自动编码器(VAE)来创建数据的紧凑潜在表示;然而,VAE对于所有设备和所有可能的私有标签保持固定。然后,我们训练一个小的生成过滤器,以扰动基于个人偏好的私人和实用信息的潜在表示。小滤波器通过利用可以在分布式设备上进行的GAN型鲁棒优化来训练。我们在三个流行的数据集上进行了实验:MNIST,UCI-Adult和CelebA,并给出了一个全面的评估,包括可视化的潜在嵌入的几何形状和估计的经验互信息,以显示我们的方法的有效性。
Distributed devices such as mobile phones can produce and store large amounts of data that can enhance machine learning models; however, this data may contain private information specific to the data owner that prevents the release of the data. We wish to reduce the correlation between user-specific private information and data while maintaining the useful information. Rather than learning a large model to achieve privatization from end to end, we introduce a decoupling of the creation of a latent representation and the privatization of data that allows user-specific privatization to occur in a distributed setting with limited computation and minimal disturbance on the utility of the data. We leverage a Variational Autoencoder (VAE) to create a compact latent representation of the data; however, the VAE remains fixed for all devices and all possible private labels. We then train a small generative filter to perturb the latent representation based on individual preferences regarding the private and utility information. The small filter is trained by utilizing a GAN-type robust optimization that can take place on a distributed device. We conduct experiments on three popular datasets: MNIST, UCI-Adult, and CelebA, and give a thorough evaluation including visualizing the geometry of the latent embeddings and estimating the empirical mutual information to show the effectiveness of our approach.