Practical Collaborative Learning for Crowdsensing in the Internet of Things with Differential Privacy

Practical Collaborative Learning for Crowdsensing in the Internet of Things with Differential Privacy
复制标题

具有差异隐私的物联网中群体感知的实用协作学习

DOI:
--
复制
发表时间:
2018
期刊:
IEEE Conference on Communications and Network Security
影响因子:
--
通讯作者:
Yanmin Gong
Yanmin Gong
中科院分区:
--
文献类型:
--
作者:
Yuanxiong Guo;Yanmin Gong

文献摘要

被引文献

相似文献

机器学习越来越多地用于为人群感知应用(如健康监测和查询建议)生成预测模型。这些模型在从不同来源收集的大量数据上训练时更准确。然而,如此大规模的数据收集带来了严重的隐私问题。照片、语音记录和位置等个人群体感知数据通常高度敏感,一旦发送给收集公司,福尔斯不受拥有这些数据的群体感知用户的控制,这可能会妨碍将所有用户数据传输到一个中心位置并使用传统机器学习方法进行训练的做法。在本文中,我们提倡一种替代方法,将数据存储在用户端,并通过在迭代过程中协调众测用户的本地训练来学习共享模型。具体来说,我们专注于正则化的经验风险最小化,并提出了一个有效的方案,基于分解,使多个crowdsensing用户共同学习一个准确的学习模型,为一个给定的学习目标,而不共享他们的私人crowdsensing数据。我们利用这样一个事实,即在许多学习任务中使用的优化问题是可分解的,并且可以通过交替方向乘法器(ADMM)以并行和分布式的方式解决。考虑到实际应用中不同用户设备的异构性,我们提出了一种异步ADMM算法来加快训练过程。我们的方案允许用户在自己的众测数据上独立训练,只共享一些更新的模型参数,而不是原始数据。此外,安全计算和分布式噪声生成新颖地集成在我们的计划,以保证在异步ADMM算法的执行共享参数的差分隐私。我们分析了隐私保证,并证明了我们的隐私保护的协作学习计划的隐私效用权衡经验的基础上,现实世界的数据。
Machine learning is increasingly used to produce predictive models for crowdsensing applications such as health monitoring and query suggestion. These models are more accurate when trained on large amount of data collected from different sources. However, such massive data collection presents serious privacy concerns. The personal crowdsensing data such as photos, voice records, and locations is often highly sensitive, and once being sent out to the collecting companies, falls out of the control of the crowdsensing users who own it. This may preclude the practice of transmitting all user data to a central location and training there using conventional machine learning approaches. In this paper, we advocate an alternative approach that leaves data stored on the user side and learns a shared model by coordinating local training of crowdsensing users in an iterative process. Specifically, we focus on regularized empirical risk minimization and propose an efficient scheme based on decomposition that enables multiple crowdsensing users to jointly learn an accurate learning model for a given learning objective without sharing their private crowdsensing data. We exploit the fact that the optimization problems used in many learning tasks are decomposable and can be solved in a parallel and distributed way by the alternating direction method of multipliers (ADMM). Considering the heterogeneity of different user devices in practice, we propose an asynchronous ADMM algorithm to speed up the training process. Our scheme lets users train independently on their own crowdsensing data and only share some updated model parameters instead of raw data. Moreover, secure computation and distributed noise generation are novelly integrated in our scheme to guarantee differential privacy of the shared parameters in the execution of the asynchronous ADMM algorithm. We analyze the privacy guarantee and demonstrate the privacy-utility trade-off of our privacy-preserving collaborative learning scheme empirically based on real-world data.