Incorporating Lexical Priors into Topic Models

Incorporating Lexical Priors into Topic Models
复制标题

DOI:
--
复制
发表时间:
2012-04
期刊:
--
影响因子:
--
通讯作者:
Jagadeesh Jagarlamudi;Hal Daumé;Raghavendra Udupa
Jagadeesh Jagarlamudi;Hal Daumé;Raghavendra Udupa
中科院分区:
其他
文献类型:
--
作者:
Jagadeesh Jagarlamudi;Hal Daumé;Raghavendra Udupa

文献摘要

被引文献

相似文献

主题模型在帮助用户理解文档语料库方面具有巨大的潜力。这种潜力受到其纯粹无监督性质的阻碍,这通常导致主题在外部任务中既不完全有意义也不有效(Chang et al.,2009年)。我们提出了一个简单而有效的方法来指导主题模型,以学习用户感兴趣的主题。我们通过提供用户认为代表语料库中潜在主题的种子词集来实现这一点。我们的模型使用这些种子来改善主题词分布(通过偏置主题来产生适当的种子词)和改善文档主题分布(通过偏置文档来选择与它们包含的种子词相关的主题)。对文档聚类任务的外部评估显示,使用种子信息时,即使在其他模型中使用种子信息天真的显着改善。
Topic models have great potential for helping users understand document corpora. This potential is stymied by their purely unsupervised nature, which often leads to topics that are neither entirely meaningful nor effective in extrinsic tasks (Chang et al., 2009). We propose a simple and effective way to guide topic models to learn topics of specific interest to a user. We achieve this by providing sets of seed words that a user believes are representative of the underlying topics in a corpus. Our model uses these seeds to improve both topic-word distributions (by biasing topics to produce appropriate seed words) and to improve document-topic distributions (by biasing documents to select topics related to the seed words they contain). Extrinsic evaluation on a document clustering task reveals a significant improvement when using seed information, even over other models that use seed information naively.