Non-readily identifiable data collaboration analysis for multiple datasets including personal information

Non-readily identifiable data collaboration analysis for multiple datasets including personal information
复制标题

DOI:
10.48550/arxiv.2208.14611
复制
发表时间:
2022-08
期刊:
Inf. Fusion
影响因子:
--
通讯作者:
A. Imakura;T. Sakurai;Yukihiko Okada;Tomoya Fujii;Teppei Sakamoto;Hiroyuki Abe
A. Imakura;T. Sakurai;Yukihiko Okada;Tomoya Fujii;Teppei Sakamoto;Hiroyuki Abe
中科院分区:
其他
文献类型:
--
作者:
A. Imakura;T. Sakurai;Yukihiko Okada;Tomoya Fujii;Teppei Sakamoto;Hiroyuki Abe

文献摘要

相似文献

多源数据融合是对多个数据源进行联合分析以获得改进信息的一种研究方法。对于多个医疗机构的数据集,数据保密和跨机构沟通至关重要。在这种情况下,通过共享降维的中间表示来进行数据协作(DC)分析,而不需要反复的跨机构通信,可能是合适的。在分析包括个人信息在内的数据时,共享数据的可识别性至关重要。在本研究中,研究了直流分析的可识别性。结果表明,共享的中间表示很容易识别为监督学习的原始数据。然后,本研究提出了一种非易于识别的DC分析,仅对包括个人信息在内的多个医疗数据集共享不易识别的数据。该方法基于随机样本排列、可解释DC分析的概念和不可重构函数的使用,解决了可识别性问题。在医学数据集的数值实验中,该方法在保持传统DC分析的高识别性能的同时,表现出不易识别性。对于医院数据集,所提出的方法在识别性能方面比仅使用本地数据集的本地分析提高了9个百分点。
Multi-source data fusion, in which multiple data sources are jointly analyzed to obtain improved information, has considerable research attention. For the datasets of multiple medical institutions, data confidentiality and cross-institutional communication are critical. In such cases, data collaboration (DC) analysis by sharing dimensionality-reduced intermediate representations without iterative cross-institutional communications may be appropriate. Identifiability of the shared data is essential when analyzing data including personal information. In this study, the identifiability of the DC analysis is investigated. The results reveals that the shared intermediate representations are readily identifiable to the original data for supervised learning. This study then proposes a non-readily identifiable DC analysis only sharing non-readily identifiable data for multiple medical datasets including personal information. The proposed method solves identifiability concerns based on a random sample permutation, the concept of interpretable DC analysis, and usage of functions that cannot be reconstructed. In numerical experiments on medical datasets, the proposed method exhibits a non-readily identifiability while maintaining a high recognition performance of the conventional DC analysis. For a hospital dataset, the proposed method exhibits a nine percentage point improvement regarding the recognition performance over the local analysis that uses only local dataset.