Federated Matrix Factorization with Privacy Guarantee

Federated Matrix Factorization with Privacy Guarantee
复制标题

DOI:
10.14778/3503585.3503598
复制
发表时间:
2021-12
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Zitao Li;Bolin Ding;Ce Zhang;Ninghui Li;Jingren Zhou
Zitao Li;Bolin Ding;Ce Zhang;Ninghui Li;Jingren Zhou
中科院分区:
其他
文献类型:
--
作者:
Zitao Li;Bolin Ding;Ce Zhang;Ninghui Li;Jingren Zhou

文献摘要

相似文献

矩阵分解(MF)是对评级矩阵中未观察到的评级进行近似,其行对应于用户,列对应于要评级的项目,并已成为推荐系统中的基本构建块。本文全面研究了不同联邦学习(FL)环境下的矩阵分解问题,其中一组参与方希望在训练中合作,但拒绝直接共享数据。我们首先提出了一个通用的算法框架,联邦矩阵分解(FMF)的各种设置,并提供了理论上的收敛保证。然后,我们系统地描述了三种不同设置的数据收集,培训和发布阶段的隐私泄露风险,并引入隐私概念,以提供端到端的隐私保护。第一种是垂直联邦学习(Vertical Federated Learning,VFL),其中多方拥有来自同一组用户的评分,但对不相交的项目集进行评分。第二种是水平联邦学习(HFL),各方获得来自不同用户集但对同一项目集的评级。第三种设置是本地联合学习(LFL),其中用户的评级仅存储在其本地设备上。我们介绍适应版本的FMF的隐私概念,保证在三个设置。特别是,一个新的私人学习技术,称为嵌入裁剪的介绍和使用在所有三个设置,以确保差异隐私。对于LFL设置,我们将联合收割机差分隐私与安全聚合相结合,以保护用户设备与服务器之间的通信,其强度类似于本地差分隐私模型,但准确性更高。我们进行实验,以证明我们的方法的有效性。
Matrix factorization (MF) approximates unobserved ratings in a rating matrix, whose rows correspond to users and columns correspond to items to be rated, and has been serving as a fundamental building block in recommendation systems. This paper comprehensively studies the problem of matrix factorization in different federated learning (FL) settings, where a set of parties want to cooperate in training but refuse to share data directly. We first propose a generic algorithmic framework for various settings of federated matrix factorization (FMF) and provide a theoretical convergence guarantee. We then systematically characterize privacy-leakage risks in data collection, training, and publishing stages for three different settings and introduce privacy notions to provide end-to-end privacy protections. The first one is vertical federated learning (VFL), where multiple parties have the ratings from the same set of users but on disjoint sets of items. The second one is horizontal federated learning (HFL), where parties have ratings from different sets of users but on the same set of items. The third setting is local federated learning (LFL), where the ratings of the users are only stored on their local devices. We introduce adapted versions of FMF with the privacy notions guaranteed in the three settings. In particular, a new private learning technique called embedding clipping is introduced and used in all the three settings to ensure differential privacy. For the LFL setting, we combine differential privacy with secure aggregation to protect the communication between user devices and the server with a strength similar to the local differential privacy model, but much better accuracy. We perform experiments to demonstrate the effectiveness of our approaches.