Discovering Relevance-Dependent Bicluster Structure from Relational Data

Discovering Relevance-Dependent Bicluster Structure from Relational Data
复制标题

DOI:
10.24963/ijcai.2017/359
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Iku Ohama;Takuya Kida;Hiroki Arimura
Iku Ohama;Takuya Kida;Hiroki Arimura
中科院分区:
其他
文献类型:
--
作者:
Iku Ohama;Takuya Kida;Hiroki Arimura

文献摘要

相似文献

在本文中,我们提出了一个统计模型的相关依赖双聚类分析关系数据。该模型将关系数据分解为具有两个特征的双簇结构:(1)簇中的每个对象都有一个相关值,该值表示对象与簇的相关程度;(2)所有簇都与至少一个密集块相关。这些功能简化了理解每个聚类的意义的任务,因为只需要检查几个高度相关的对象。我们引入相关性依赖的伯努利分布(R-BD)作为相关性依赖的二元矩阵的先验,并提出了新的相关性依赖的无限双聚类(R-IB)模型,该模型自动估计聚类数。后验推理可以有效地使用一个崩溃的吉布斯采样器,因为R-IB模型的参数可以完全边缘化。实验结果表明,R-IB提取更多的本质双集群结构,更好的计算效率比传统的模型。我们进一步观察到,RIB获得的双聚类结果有助于解释每个聚类的含义。
In this paper, we propose a statistical model for relevance-dependent biclustering to analyze relational data. The proposed model factorizes relational data into bicluster structure with two features: (1) each object in a cluster has a relevance value, which indicates how strongly the object relates to the cluster and (2) all clusters are related to at least one dense block. These features simplify the task of understanding the meaning of each cluster because only a few highly relevant objects need to be inspected. We introduced the RelevanceDependent Bernoulli Distribution (R-BD) as a prior for relevance-dependent binary matrices and proposed the novel Relevance-Dependent Infinite Biclustering (R-IB) model, which automatically estimates the number of clusters. Posterior inference can be performed efficiently using a collapsed Gibbs sampler because the parameters of the R-IB model can be fully marginalized out. Experimental results show that the R-IB extracts more essential bicluster structure with better computational efficiency than conventional models. We further observed that the biclustering results obtained by RIB facilitate interpretation of the meaning of each cluster.