Towards Task-free Privacy-preserving Data Collection

Towards Task-free Privacy-preserving Data Collection
复制标题

实现无任务的隐私保护数据收集

DOI:
10.23919/jcc.2022.07.023
复制
发表时间:
2022-07
影响因子:
4.1
通讯作者:
Huajie Shao
Huajie Shao
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhibo Wang;Wei Yuan;Xiaoyi Pang;Jingxin Li;Huajie Shao

文献摘要

相似文献

随着物联网(IoT)的快速发展和嵌入式设备的激增,收集了大量的个人数据,然而,这些数据可能携带大量关于用户不想共享的属性的私有信息。已经提出了许多隐私保护方法,以防止隐私泄漏,通过扰动原始数据或在本地设备上提取面向任务的特征。不幸的是,它们在应用于其他任务时会遭受严重的隐私泄漏和准确性下降,因为它们是为预定义的任务设计和优化的。在本文中,我们提出了一种新的无任务的隐私保护数据收集方法,通过对抗表示学习,称为TF-ARL,保护用户指定的私有属性,同时保持未知的下游任务的数据效用。为此,我们首先提出了一种隐私对抗学习机制(PAL),通过优化特征提取器来保护私有属性,以最大限度地提高对手对私有属性的预测不确定性,然后设计了一种条件解码机制(ConDec),通过最小化来自净化特征的条件重建误差来维护下游任务的数据效用。通过PAL和ConDec的联合学习,我们可以学习一个隐私感知的特征提取器,其中净化后的特征保持了除隐私之外的区别性信息。在真实数据集上的大量实验结果证明了TF-ARL的有效性。
With the rapid developments of Internet of Things (IoT) and proliferation of embedded devices, large volume of personal data are collected, which however, might carry massive private information about attributes that users do not want to share. Many privacy-preserving methods have been proposed to prevent privacy leakage by perturbing raw data or extracting task-oriented features at local devices. Unfortunately, they would suffer from significant privacy leakage and accuracy drop when applied to other tasks as they are designed and optimized for predefined tasks. In this paper, we propose a novel task-free privacy-preserving data collection method via adversarial representation learning, called TF-ARL, to protect private attributes specified by users while maintaining data utility for unknown downstream tasks. To this end, we first propose a privacy adversarial learning mechanism (PAL) to protect private attributes by optimizing the feature extractor to maximize the adversary's prediction uncertainty on private attributes, and then design a conditional decoding mechanism (ConDec) to maintain data utility for downstream tasks by minimizing the conditional reconstruction error from the sanitized features. With the joint learning of PAL and ConDec, we can learn a privacy-aware feature extractor where the sanitized features maintain the discriminative information except privacy. Extensive experimental results on real-world datasets demonstrate the effectiveness of TF-ARL.