Automatic Labelling of Topic Models Learned from Twitter by Summarisation

Automatic Labelling of Topic Models Learned from Twitter by Summarisation
复制标题

DOI:
10.3115/v1/p14-2101
复制
发表时间:
2014-06
期刊:
--
影响因子:
--
通讯作者:
A. Cano;Yulan He;Ruifeng Xu
A. Cano;Yulan He;Ruifeng Xu
中科院分区:
其他
文献类型:
--
作者:
A. Cano;Yulan He;Ruifeng Xu

文献摘要

被引文献

相似文献

潜在的主题模型,如潜在的狄利克雷分配(LDA)是隐藏的主题结构的结果,提供了进一步的见解的数据。然而,对源自社交媒体的此类主题的自动标记提出了新的挑战,因为主题可能会隐藏在真实的世界中发生的新事件。依赖于外部知识源的现有自动主题标记方法在这里变得不太适用,因为所提取的主题的相关文章/概念可能不存在于外部源中。在本文中,我们建议解决的问题,自动标记的潜在主题从Twitter的总结问题。我们介绍了一个框架,应用摘要算法生成主题标签。这些算法独立于外部源,并且仅依赖于与潜在主题相关的文档中的主导术语的识别。我们比较现有的最先进的汇总算法的效率。我们的研究结果表明,总结算法生成更好的主题标签,捕捉事件相关的上下文相比,前n项LDA返回。
Latent topics derived by topic models such as Latent Dirichlet Allocation (LDA) are the result of hidden thematic structures which provide further insights into the data. The automatic labelling of such topics derived from social media poses however new challenges since topics may characterise novel events happening in the real world. Existing automatic topic labelling approaches which depend on external knowledge sources become less applicable here since relevant articles/concepts of the extracted topics may not exist in external sources. In this paper we propose to address the problem of automatic labelling of latent topics learned from Twitter as a summarisation problem. We introduce a framework which apply summarisation algorithms to generate topic labels. These algorithms are independent of external sources and only rely on the identification of dominant terms in documents related to the latent topic. We compare the efficiency of existing state of the art summarisation algorithms. Our results suggest that summarisation algorithms generate better topic labels which capture event-related context compared to the top-n terms returned by LDA.