BTM: Topic Modeling over Short Texts

BTM: Topic Modeling over Short Texts
复制标题

BTM:短文本主题建模

DOI:
10.1109/tkde.2014.2313872
复制
发表时间:
2014-12-01
影响因子:
8.9
通讯作者:
Guo, Jiafeng
Guo, Jiafeng
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cheng, Xueqi;Yan, Xiaohui;Guo, Jiafeng

文献摘要

被引文献

相似文献

短文本在当今的网络上很受欢迎,尤其是随着社交媒体的出现。对许多内容分析任务来说,从大规模短文本中推断主题是一项关键而又具有挑战性的任务。传统的主题模型,如潜在狄利let分配(LDA)和概率潜在语义分析(PLSA),通过将每个文档建模为主题的混合,从文档级单词共现中学习主题,其推理受到短文本中单词共现模式的稀疏性的影响。本文提出了一种新的短文本主题建模方法——双词主题模型(BTM)。BTM通过对语料库中词共现模式(即双词)的生成直接建模来学习主题,利用丰富的语料库级信息进行有效的推理。为了处理大规模的短文本数据,我们进一步引入了两种在线的BTM算法来实现高效的主题学习。在真实单词短文本集合上的实验表明,BTM可以发现更突出和连贯的主题,并且显著优于最先进的基线。我们还展示了两种在线BTM算法在时间效率和主题学习方面的令人满意的性能。
Short texts are popular on today's web, especially with the emergence of social media. Inferring topics from large scale short texts becomes a critical but challenging task for many content analysis tasks. Conventional topic models such as latent Dirichlet allocation (LDA) and probabilistic latent semantic analysis (PLSA) learn topics from document-level word co-occurrences by modeling each document as a mixture of topics, whose inference suffers from the sparsity of word co-occurrence patterns in short texts. In this paper, we propose a novel way for short text topic modeling, referred as biterm topic model (BTM). BTM learns topics by directly modeling the generation of word co-occurrence patterns (i.e., biterms) in the corpus, making the inference effective with the rich corpus-level information. To cope with large scale short text data, we further introduce two online algorithms for BTM for efficient topic learning. Experiments on real-word short text collections show that BTM can discover more prominent and coherent topics, and significantly outperform the state-of-the-art baselines. We also demonstrate the appealing performance of the two online BTM algorithms on both time efficiency and topic learning.