Privacy-Preserving Data Sharing by Integrating Perturbed Distance Matrices

Privacy-Preserving Data Sharing by Integrating Perturbed Distance Matrices
复制标题

DOI:
10.1007/s42979-020-00127-w
复制
发表时间:
2020-04
期刊:
SN Computer Science
影响因子:
--
通讯作者:
Hanten Chang;H. Ando
Hanten Chang;H. Ando
中科院分区:
其他
文献类型:
--
作者:
Hanten Chang;H. Ando

文献摘要

相似文献

收集大量数据有利于机器学习生成偏差较小的模型。在许多情况下,类似的数据分散在各组织之间,由于涉及隐私和成本的问题,很难整合这些数据。在不交付原始数据的情况下集成这些分布式数据会产生数据协作的概念,它以安全的方式组合不同组织持有的数据。我们提出了一种方法,其中使用组织之间的公共数据获得的原始数据的距离矩阵被共享,以学习原始数据的邻居信息。具体而言,所提出的方法鲁棒地集成分布式数据,这是连接的原始数据的质量一样好,在每个组织中的数据量小,数据偏差大的情况下。此外,该方法适用于受噪声污染的数据。为了证明所提出的方法的有效性,我们进行了一个开放的生物数据分为几块的分类任务,发现划分数据的分类结果是精确的,当所有的数据都可用。最后,我们证明了该方法对噪声的鲁棒性提高了原始数据的匿名性。
Collecting large amounts of data is beneficial in machine learning to generate models that are less biased. There are many cases in which pieces of similar data are distributed among organizations, and it is difficult to integrate these data owing to issues involving privacy and cost. Integrating these distributed data without delivering the original data leads to the concept of data collaboration, which combines data held by different organizations in a secure manner. We propose a method in which a distance matrix of the original data obtained using common data among organizations is shared to learn neighbor information of the original data. Specifically, the proposed method robustly integrates distributed data, which is of as good quality as connected raw data, in cases where the amount of data in each organization is small and the data bias is large. In addition, the proposed method is applicable to data contaminated by noise. To demonstrate the effectiveness of the proposed method, we performed a classification task on open biological data divided into several pieces and found that the classification results for divided data were as precise as when all data were available. Finally, we show that the robustness of the method against noise improves the anonymity of the original data as a by-product.