Privacy-Preserving Multiple Linear Regression of Vertically Partitioned Real Medical Datasets

Privacy-Preserving Multiple Linear Regression of Vertically Partitioned Real Medical Datasets
复制标题

DOI:
10.1109/aina.2017.52
复制
发表时间:
2017-03
期刊:
2017 IEEE 31st International Conference on Advanced Information Networking and Applications (AINA)
影响因子:
--
通讯作者:
Hiroaki Kikuchi;Chika Hamanaga;H. Yasunaga;H. Matsui;H. Hashimoto
Hiroaki Kikuchi;Chika Hamanaga;H. Yasunaga;H. Matsui;H. Hashimoto
中科院分区:
其他
文献类型:
--
作者:
Hiroaki Kikuchi;Chika Hamanaga;H. Yasunaga;H. Matsui;H. Hashimoto

文献摘要

被引文献

相似文献

本文研究了隐私保护数据挖掘在流行病学研究中的可行性。至于数据挖掘算法,我们专注于线性多元回归,可以用来识别许多可能变量中最重要的因素,例如许多疾病的历史。我们尝试从与病人和疾病信息相关的分布式数据集中识别线性模型来估计住院时间。本文利用与脑卒中相关的真实的医学数据集进行了实验,并尝试应用年龄、性别、医学量表、日本昏迷量表和改良的兰金量表。本文的贡献包括:(1)提出了一个实用的垂直分区线性多元回归隐私保护协议;(2)利用分布在两个当事方的真实的医疗数据集,证明了所提出的系统的可行性,和当地政府谁知道的住所,甚至在病人离开医院。(3)展示了PPDM系统的准确性和性能,该系统允许我们用任意数量的预测器来估计预期的处理时间。
This paper studies the feasibility of privacy-preservingdata mining in epidemiological study. As for the data-miningalgorithm, we focus to a linear multiple regression thatcan be used to identify the most significant factorsamong many possible variables, such as the historyof many diseases. We try to identify the linear model to estimate a lengthof hospital stay from distributed dataset related tothe patient and the disease information. In this paper, we have done experiment usingthe real medical dataset related to stroke andattempt to apply multiple regression with sixpredictors of age, sex, the medical scales, e.g., Japan Coma Scale, and the modified Rankin Scale. Our contributions of this paper include(1) to propose a practical privacy-preserving protocols for linear multiple regressionwith vertically partitioned datasets, and(2) to show the feasibility of the proposed system usingthe real medical dataset distributed into two parties, the hospital who knows the technical details of diseasesduring the patients are in the hospital, and the local government who knows the residence even afterthe patients left hospital. (3) to show the accuracy and the performance of thePPDM system which allows us to estimate the expectedprocessing time with arbitrary number of predictors.