Privacy-Preserving Distributed Linear Regression on High-Dimensional Data

Privacy-Preserving Distributed Linear Regression on High-Dimensional Data
复制标题

DOI:
10.1515/popets-2017-0053
复制
发表时间:
2017-10
影响因子:
--
通讯作者:
Adrià Gascón;Phillipp Schoppmann;Borja Balle;Mariana Raykova;Jack Doerner;Samee Zahur;David Evans
Adrià Gascón;Phillipp Schoppmann;Borja Balle;Mariana Raykova;Jack Doerner;Samee Zahur;David Evans
中科院分区:
--
文献类型:
--
作者:
Adrià Gascón;Phillipp Schoppmann;Borja Balle;Mariana Raykova;Jack Doerner;Samee Zahur;David Evans

文献摘要

被引文献

相似文献

摘要:我们提出了一种隐私保护协议,用于计算线性回归模型,在训练数据集垂直分布在几个方面的设置。我们的主要贡献是一个混合多方计算协议,结合姚的乱码电路与定制的协议计算内积。与许多机器学习任务一样,构建线性回归模型涉及求解线性方程组。我们对安全执行此任务的不同技术进行了全面的评估和比较,包括新的共轭梯度下降(CGD)算法。该算法适用于安全计算,因为它使用了真实的数的有效定点表示,同时保持了与使用浮点数的经典解决方案可获得的精度和收敛速度相当的精度和收敛速度。我们的技术改进了Nikolaenko等人的技术。的方法用于隐私保护岭回归(S&P 2013),并可用作其他分析的构建块。我们实现了一个完整的系统,并证明了我们的方法是高度可扩展的,解决数据分析问题,在不到一个小时的总运行时间与一百万条记录和一百个功能。
Abstract We propose privacy-preserving protocols for computing linear regression models, in the setting where the training dataset is vertically distributed among several parties. Our main contribution is a hybrid multi-party computation protocol that combines Yao’s garbled circuits with tailored protocols for computing inner products. Like many machine learning tasks, building a linear regression model involves solving a system of linear equations. We conduct a comprehensive evaluation and comparison of different techniques for securely performing this task, including a new Conjugate Gradient Descent (CGD) algorithm. This algorithm is suitable for secure computation because it uses an efficient fixed-point representation of real numbers while maintaining accuracy and convergence rates comparable to what can be obtained with a classical solution using floating point numbers. Our technique improves on Nikolaenko et al.’s method for privacy-preserving ridge regression (S&P 2013), and can be used as a building block in other analyses. We implement a complete system and demonstrate that our approach is highly scalable, solving data analysis problems with one million records and one hundred features in less than one hour of total running time.