Bayesian Nonparametric Relational Topic Model through Dependent Gamma Processes

Bayesian Nonparametric Relational Topic Model through Dependent Gamma Processes
复制标题

通过相关伽马过程的贝叶斯非参数关系主题模型

DOI:
10.1109/tkde.2016.2636182
复制
发表时间:
2017-07
影响因子:
8.9
通讯作者:
Xiangfeng Luo
Xiangfeng Luo
中科院分区:
计算机科学2区
文献类型:
--
作者:
Junyu Xuan;Jie Lu;Guangquan Zhang;Richard Yi Da Xu;Xiangfeng Luo

文献摘要

参考文献

被引文献

相似文献

传统的关系主题模型为发现文档网络中的隐藏主题提供了一种成功的方法。许多理论和实际任务,如降维,文档聚类和链接预测,可以受益于这些揭示的知识。然而,现有的关系主题模型是基于一个假设,即隐藏主题的数量是已知的先验,这是不切实际的,在许多现实世界的应用。因此,为了放松这一假设,我们提出了一个非参数的关系主题模型,使用随机过程,而不是固定维的概率分布在本文中。具体来说,每个文档都被分配了一个Gamma过程,它代表了这个文档的主题兴趣。虽然这种方法提供了一种优雅的解决方案,但在对典型文档网络的固有网络结构进行数学建模时,它带来了额外的挑战,即,两个空间上更接近的文档往往具有更相似的主题。此外,我们要求所有文档共享主题。为了解决这些挑战,我们使用一个二次抽样策略,每个文件分配一个不同的Gamma过程从全局Gamma过程,和文件的二次抽样概率分配与马尔可夫随机场约束,继承了文件的网络结构。通过设计的后验推理算法,可以同时发现隐藏主题及其数量。在合成和真实网络数据集上的实验结果表明,该算法能够学习隐藏主题,更重要的是学习主题的数量。
Traditional relational topic models provide a successful way to discover the hidden topics from a document network. Many theoretical and practical tasks, such as dimensional reduction, document clustering, and link prediction, could benefit from this revealed knowledge. However, existing relational topic models are based on an assumption that the number of hidden topics is known a priori, which is impractical in many real-world applications. Therefore, in order to relax this assumption, we propose a nonparametric relational topic model using stochastic processes instead of fixed-dimensional probability distributions in this paper. Specifically, each document is assigned a Gamma process, which represents the topic interest of this document. Although this method provides an elegant solution, it brings additional challenges when mathematically modeling the inherent network structure of typical document network, i.e., two spatially closer documents tend to have more similar topics. Furthermore, we require that the topics are shared by all the documents. In order to resolve these challenges, we use a subsampling strategy to assign each document a different Gamma process from the global Gamma process, and the subsampling probabilities of documents are assigned with a Markov Random Field constraint that inherits the document network structure. Through the designed posterior inference algorithm, we can discover the hidden topics and its number simultaneously. Experimental results on both synthetic and real-world network datasets demonstrate the capabilities of learning the hidden topics and, more importantly, the number of topics.
DOI: 10.1007/978-3-540-30115-8_22
发表时间: 2004-09
期刊: --
影响因子: --
作者:
Bryan Klimt;Yiming Yang
通讯作者: Bryan Klimt;Yiming Yang
DOI: --
发表时间: 2012-06
期刊: ArXiv
影响因子: --
作者:
Yingjian Wang;L. Carin
通讯作者: Yingjian Wang;L. Carin
DOI: --
发表时间: 2009-04
期刊: --
影响因子: --
作者:
Jonathan D. Chang;D. Blei
通讯作者: Jonathan D. Chang;D. Blei
DOI: 10.5555/1390681.1442798
发表时间: 2007-05
期刊: Journal of machine learning research : JMLR
影响因子: --
作者:
E. Airoldi;D. Blei;S. Fienberg;E. Xing
通讯作者: E. Airoldi;D. Blei;S. Fienberg;E. Xing
DOI: 10.1214/10-aoas395
发表时间: 2011-06-01
影响因子: 1.8
作者:
Fox, Emily B.;Sudderth, Erik B.;Willsky, Alan S.
通讯作者: Willsky, Alan S.