Privacy-Preserving Collaborative Prediction using Random Forests

Privacy-Preserving Collaborative Prediction using Random Forests
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science
影响因子:
--
通讯作者:
Irene Giacomelli;S. Jha;Ross Kleiman;David Page;Kyonghwan Yoon
Irene Giacomelli;S. Jha;Ross Kleiman;David Page;Kyonghwan Yoon
中科院分区:
其他
文献类型:
--
作者:
Irene Giacomelli;S. Jha;Ross Kleiman;David Page;Kyonghwan Yoon

文献摘要

被引文献

相似文献

我们研究了集成方法的隐私保护机器学习(PPML)问题,重点研究了随机森林。在协同分析中,PPML试图解决数据共享需求与隐私之间的冲突。这在隐私敏感的应用中尤其重要,例如从不同诊所的电子病历数据中学习临床决策支持的预测模型,其中每个诊所都对其患者的隐私负责。我们提出了一种集成方法的新方法:每个实体从自己的数据中学习一个模型,然后当客户端要求预测一个新的私有实例时,来自所有局部训练模型的答案被用来计算预测,这样就不会透露额外的信息。我们在随机森林中实现了这种方法,并通过对包括实际EHR数据在内的真实数据集的实验证明了它的高效率和潜在的准确性优势。
We study the problem of privacy-preserving machine learning (PPML) for ensemble methods, focusing our effort on random forests. In collaborative analysis, PPML attempts to solve the conflict between the need for data sharing and privacy. This is especially important in privacy sensitive applications such as learning predictive models for clinical decision support from EHR data from different clinics, where each clinic has a responsibility for its patients' privacy. We propose a new approach for ensemble methods: each entity learns a model, from its own data, and then when a client asks the prediction for a new private instance, the answers from all the locally trained models are used to compute the prediction in such a way that no extra information is revealed. We implement this approach for random forests and we demonstrate its high efficiency and potential accuracy benefit via experiments on real-world datasets, including actual EHR data.