Privacy-preserving logistic regression training.

Privacy-preserving logistic regression training.
复制标题

DOI:
10.1186/s12920-018-0398-y
复制
发表时间:
2018-10-11
影响因子:
2.7
通讯作者:
Vercauteren F
Vercauteren F
中科院分区:
医学3区
文献类型:
--
作者:
Bonte C;Vercauteren F

文献摘要

参考文献

被引文献

相似文献

逻辑回归是机器学习中用于构建分类模型的流行技术。由于这种模型的构建是基于对大型数据集的计算,因此将此计算外包给云服务是一个有吸引力的想法。输入数据的隐私敏感性质要求在外包之前采取适当的隐私保护措施。同态加密使人们能够直接对加密数据进行计算,而无需解密,并可用于减轻使用云服务引起的隐私问题。在本文中,我们提出了一种算法(及其实现)来训练同态加密数据集上的逻辑回归模型。我们的算法的核心包括一个新的迭代方法,可以看作是一个简化形式的固定海森方法,但具有低得多的乘法复杂度。我们在两个有趣的真实的生活应用中测试了新方法:第一个应用是在医学中,构建一个模型来预测患者患癌症的概率,给定基因组数据作为输入;第二个应用是在金融中,该模型预测信用卡交易的概率是欺诈性的。该方法为两个应用程序产生准确的结果,与在纯文本数据上运行标准算法相当。本文介绍了一种新的简单迭代算法来训练逻辑回归模型,该模型适用于同态加密数据集。该算法可以作为一种隐私保护技术来建立一个二进制分类模型,并可以应用于广泛的问题,可以用逻辑回归建模。我们的实现结果表明,我们的方法可以处理逻辑回归训练中使用的大数据集。
Logistic regression is a popular technique used in machine learning to construct classification models. Since the construction of such models is based on computing with large datasets, it is an appealing idea to outsource this computation to a cloud service. The privacy-sensitive nature of the input data requires appropriate privacy preserving measures before outsourcing it. Homomorphic encryption enables one to compute on encrypted data directly, without decryption and can be used to mitigate the privacy concerns raised by using a cloud service. In this paper, we propose an algorithm (and its implementation) to train a logistic regression model on a homomorphically encrypted dataset. The core of our algorithm consists of a new iterative method that can be seen as a simplified form of the fixed Hessian method, but with a much lower multiplicative complexity. We test the new method on two interesting real life applications: the first application is in medicine and constructs a model to predict the probability for a patient to have cancer, given genomic data as input; the second application is in finance and the model predicts the probability of a credit card transaction to be fraudulent. The method produces accurate results for both applications, comparable to running standard algorithms on plaintext data. This article introduces a new simple iterative algorithm to train a logistic regression model that is tailored to be applied on a homomorphically encrypted dataset. This algorithm can be used as a privacy-preserving technique to build a binary classification model and can be applied in a wide range of problems that can be modelled with logistic regression. Our implementation results show that our method can handle the large datasets used in logistic regression training.
DOI: 10.1007/bf00048682
发表时间: 1992-03-01
影响因子: 1
作者:
BOHNING, D
通讯作者: BOHNING, D
DOI: 10.1016/j.jbi.2014.04.003
发表时间: 2014-08-01
影响因子: 4.5
作者:
Bos, Joppe W.;Lauter, Kristin;Naehrig, Michael
通讯作者: Naehrig, Michael
DOI: 10.1007/bf00049423
发表时间: 1988-01-01
影响因子: 1
作者:
BOHNING, D;LINDSAY, BG
通讯作者: LINDSAY, BG
DOI: 10.1145/2535925
发表时间: 2013-11-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Lyubashevsky, Vadim;Peikert, Chris;Regev, Oded
通讯作者: Regev, Oded