OpBoost: A Vertical Federated Tree Boosting Framework Based on Order-Preserving Desensitization

OpBoost: A Vertical Federated Tree Boosting Framework Based on Order-Preserving Desensitization
复制标题

DOI:
10.14778/3565816.3565823
复制
发表时间:
2022-10
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Xiaochen Li;Yuke Hu;Weiran Liu;Hanwen Feng;Li Peng;Yuan Hong;Kui Ren;Zhan Qin
Xiaochen Li;Yuke Hu;Weiran Liu;Hanwen Feng;Li Peng;Yuan Hong;Kui Ren;Zhan Qin
中科院分区:
其他
文献类型:
--
作者:
Xiaochen Li;Yuke Hu;Weiran Liu;Hanwen Feng;Li Peng;Yuan Hong;Kui Ren;Zhan Qin

文献摘要

相似文献

垂直联合学习(FL)是一种新的学习范式,它允许具有相同数据样本非重叠属性的用户联合训练模型,而不直接共享原始数据。然而,最近的研究表明,这仍然不足以防止训练过程或训练模型中的隐私泄露。本文重点研究了垂直FL下的隐私保护树提升算法。现有的基于密码学的解决方案计算量大,通信开销大,容易受到推理攻击。虽然基于局部差分隐私(LDP)的解决方案解决了上述问题,但它导致训练模型的准确率较低。针对目前广泛应用的满足垂直FL环境下差异隐私的树提升算法的准确性问题进行了研究。具体地说,我们引入了一个名为OpBoost的框架。设计了三种满足LDP的变种--基于距离的LDP(DLDP)的保序去敏算法,用于训练数据的去敏感。特别是,我们对dLDP的定义进行了优化,并研究了有效的采样分布,以进一步提高所提出算法的精度和效率。所提出的算法在大距离配对的私密性和不敏感值的效用之间提供了折衷。综合评价表明,在合理的设置下,OpBoost在训练模型的预测精度上优于现有的LDP方法。我们的代码是开源的。
Vertical Federated Learning (FL) is a new paradigm that enables users with non-overlapping attributes of the same data samples to jointly train a model without directly sharing the raw data. Nevertheless, recent works show that it's still not sufficient to prevent privacy leakage from the training process or the trained model. This paper focuses on studying the privacy-preserving tree boosting algorithms under the vertical FL. The existing solutions based on cryptography involve heavy computation and communication overhead and are vulnerable to inference attacks. Although the solution based on Local Differential Privacy (LDP) addresses the above problems, it leads to the low accuracy of the trained model. This paper explores to improve the accuracy of the widely deployed tree boosting algorithms satisfying differential privacy under vertical FL. Specifically, we introduce a framework called OpBoost. Three order-preserving desensitization algorithms satisfying a variant of LDP called distance-based LDP (dLDP) are designed to desensitize the training data. In particular, we optimize the dLDP definition and study efficient sampling distributions to further improve the accuracy and efficiency of the proposed algorithms. The proposed algorithms provide a trade-off between the privacy of pairs with large distance and the utility of desensitized values. Comprehensive evaluations show that OpBoost has a better performance on prediction accuracy of trained models compared with existing LDP approaches on reasonable settings. Our code is open source.