Automatic labeling of multinomial topic models

Automatic labeling of multinomial topic models
复制标题

DOI:
10.1145/1281192.1281246
复制
发表时间:
2007-08
影响因子:
8.7
通讯作者:
Qiaozhu Mei;Xuehua Shen;ChengXiang Zhai
Qiaozhu Mei;Xuehua Shen;ChengXiang Zhai
中科院分区:
农林科学1区
文献类型:
--
作者:
Qiaozhu Mei;Xuehua Shen;ChengXiang Zhai

文献摘要

被引文献

相似文献

单词上的多项分布经常用于对文本集合中的主题进行建模。在将所有此类主题模型应用于任何文本挖掘问题时,一个常见的主要挑战是准确地标记多项式主题模型,以便用户可以解释发现的主题。到目前为止,这些标签都是以主观的方式手工生成的。本文提出了一种客观地自动标注多项式主题模型的概率方法。我们将这个标注问题视为一个优化问题,涉及最小化词分布之间的Kullback-Leibler散度和最大化标签和主题模型之间的互信息。在两个不同体裁的文本数据集上进行了用户研究实验。结果表明,所提出的标记方法能够有效地生成对发现的主题模型有意义和有用的标签。我们的方法是通用的,可以应用于标记通过各种主题模型(如PLSA, LDA及其变体)学习的主题。
Multinomial distributions over words are frequently used to model topics in text collections. A common, major challenge in applying all such topic models to any text mining problem is to label a multinomial topic model accurately so that a user can interpret the discovered topic. So far, such labels have been generated manually in a subjective way. In this paper, we propose probabilistic approaches to automatically labeling multinomial topic models in an objective way. We cast this labeling problem as an optimization problem involving minimizing Kullback-Leibler divergence between word distributions and maximizing mutual information between a label and a topic model. Experiments with user study have been done on two text data sets with different genres.The results show that the proposed labeling methods are quite effective to generate labels that are meaningful and useful for interpreting the discovered topic models. Our methods are general and can be applied to labeling topics learned through all kinds of topic models such as PLSA, LDA, and their variations.