A Collaborative Framework for Privacy Preserving Fuzzy Co-clustering of Vertically Distributed Cooccurrence Matrices

A Collaborative Framework for Privacy Preserving Fuzzy Co-clustering of Vertically Distributed Cooccurrence Matrices
复制标题

垂直分布共生矩阵的隐私保护模糊共聚类协作框架

DOI:
10.1155/2015/729072
复制
发表时间:
2015
影响因子:
1.3
通讯作者:
A. Notsu
A. Notsu
中科院分区:
--
文献类型:
--
作者:
K. Honda;T. Oda;D. Tanaka;A. Notsu

文献摘要

相似文献

在许多现实世界的数据分析任务中,我们期望通过利用存储在不同组织(如合作小组、国家机关和盟国)中的多个数据库来获得更多有用的知识。然而,在许多这样的组织中,尽管他们相信协作分析的优势,但由于隐私和安全问题,他们经常对发布数据库犹豫不决。本文提出了一种利用垂直分割的协同矩阵进行模糊共簇结构估计的协作框架,该框架将对象和项目之间的协同信息分别存储在多个站点中。为了在不担心信息泄露的情况下利用这些分布式数据集,将隐私保护过程引入到分类多元数据的模糊聚类中。保留协同矩阵的每个元素,只有对象成员关系由多个站点共享,并且通过迭代聚类过程揭示其(隐式)联合共簇结构。几个实验结果表明,协同分析有助于揭示单独矩阵的整体内在共簇结构,而不是单个位点分析。这个新框架使许多私人和公共组织能够共享公共数据结构知识,而不必担心信息泄露。
In many real world data analysis tasks, it is expected that we can get much more useful knowledge by utilizing multiple databases stored in different organizations, such as cooperation groups, state organs, and allied countries. However, in many such organizations, they often hesitate to publish their databases because of privacy and security issues although they believe the advantages of collaborative analysis. This paper proposes a novel collaborative framework for utilizing vertically partitioned cooccurrence matrices in fuzzy co‐cluster structure estimation, in which cooccurrence information among objects and items is separately stored in several sites. In order to utilize such distributed data sets without fear of information leaks, a privacy preserving procedure is introduced to fuzzy clustering for categorical multivariate data (FCCM). Withholding each element of cooccurrence matrices, only object memberships are shared by multiple sites and their (implicit) joint co‐cluster structures are revealed through an iterative clustering process. Several experimental results demonstrate that collaborative analysis can contribute to revealing global intrinsic co‐cluster structures of separate matrices rather than individual site‐wise analysis. The novel framework makes it possible for many private and public organizations to share common data structural knowledge without fear of information leaks.