Differentially Private Random Forest with High Utility

Differentially Private Random Forest with High Utility
复制标题

DOI:
10.1109/icdm.2015.76
复制
发表时间:
2015-11
期刊:
2015 IEEE International Conference on Data Mining
影响因子:
--
通讯作者:
Santu Rana;S. Gupta;S. Venkatesh
Santu Rana;S. Gupta;S. Venkatesh
中科院分区:
其他
文献类型:
--
作者:
Santu Rana;S. Gupta;S. Venkatesh

文献摘要

被引文献

相似文献

保护隐私的数据挖掘已成为敏感数据和个人数据领域的研究热点。例如,高度敏感的医疗或财务记录数字存储库为风险预测和决策提供了巨大的价值。然而,从这些存储库派生的预测模型应该严格维护个人隐私。在差分隐私框架下提出了一种新的随机森林算法。与以往严格遵循差分隐私并保持完整数据分布在一个数据实例中近似不变的工作不同,我们只保持必要的统计量(例如估计的方差)不变。这种放松会显著提高效用。为了实现我们的方法,我们提出了一种新的差分私有决策树归纳算法,并使用它们来创建决策树集合。我们还提出了可行的对手模型,在已知所有其他数据的情况下推断未知数据的属性和类标签。在这些对手模型下,我们推导出在保持隐私的情况下集合中允许的最大树数的界限。我们专注于二元分类问题,并在四个现实世界的数据集上演示了我们的方法。与现有的隐私保护方法相比,我们实现了更高的实用性。
Privacy-preserving data mining has become an active focus of the research community in the domains where data are sensitive and personal in nature. For example, highly sensitive digital repositories of medical or financial records offer enormous values for risk prediction and decision making. However, prediction models derived from such repositories should maintain strict privacy of individuals. We propose a novel random forest algorithm under the framework of differential privacy. Unlike previous works that strictly follow differential privacy and keep the complete data distribution approximately invariant to change in one data instance, we only keep the necessary statistics (e.g. variance of the estimate) invariant. This relaxation results in significantly higher utility. To realize our approach, we propose a novel differentially private decision tree induction algorithm and use them to create an ensemble of decision trees. We also propose feasible adversary models to infer about the attribute and class label of unknown data in presence of the knowledge of all other data. Under these adversary models, we derive bounds on the maximum number of trees that are allowed in the ensemble while maintaining privacy. We focus on binary classification problem and demonstrate our approach on four real-world datasets. Compared to the existing privacy preserving approaches we achieve significantly higher utility.