Anchor-Free Correlated Topic Modeling: Identifiability and Algorithm

Anchor-Free Correlated Topic Modeling: Identifiability and Algorithm
复制标题

DOI:
--
复制
发表时间:
2016-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Kejun Huang;Xiao Fu;N. Sidiropoulos
Kejun Huang;Xiao Fu;N. Sidiropoulos
中科院分区:
其他
文献类型:
--
作者:
Kejun Huang;Xiao Fu;N. Sidiropoulos

文献摘要

被引文献

相似文献

在主题建模中,已经在存在锚词的前提下开发了许多保证主题的可识别性的算法--即,仅出现在一个主题中的单词(概率为正)。后续的工作则采用了语料库的三阶或更高阶统计量来放松对锚词的假设。然而,高阶统计量的可靠估计难以获得,而且根据这些模型确定主题取决于主题的不相关性,这可能是不现实的。本文回顾了基于二阶矩的主题建模,提出了一个无锚点的主题挖掘框架。与锚词假设相比,所提出的方法保证了在更温和的条件下识别主题,从而在实践中表现出更好的鲁棒性。相应的算法只涉及一个特征分解和几个小的线性规划。这使得它易于实现和扩展到非常大的问题实例。使用TDT 2和路透社-21578语料库的实验表明,所提出的无锚方法表现出非常有利的性能(使用一致性,相似性计数,和聚类精度度量)相比,现有技术。
In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the anchor word assumption. Reliable estimates of higher-order statistics are hard to obtain, however, and the identification of topics under those models hinges on uncorrelatedness of the topics, which can be unrealistic. This paper revisits topic modeling based on second-order moments, and proposes an anchor-free topic mining framework. The proposed approach guarantees the identification of the topics under a much milder condition compared to the anchor-word assumption, thereby exhibiting much better robustness in practice. The associated algorithm only involves one eigen-decomposition and a few small linear programs. This makes it easy to implement and scale up to very large problem instances. Experiments using the TDT2 and Reuters-21578 corpus demonstrate that the proposed anchor-free approach exhibits very favorable performance (measured using coherence, similarity count, and clustering accuracy metrics) compared to the prior art.