A T EXT MINING RESEARCH BASED ON LDA T OPIC MODELLING

A T EXT MINING RESEARCH BASED ON LDA T OPIC MODELLING
复制标题

基于LDA主题建模的文本挖掘研究

DOI:
10.5121/csit.2016.60616
复制
发表时间:
2016
期刊:
ArXiv
影响因子:
--
通讯作者:
Haiyi Zhang
Haiyi Zhang
中科院分区:
--
文献类型:
--
作者:
Zhou Tong;Haiyi Zhang

文献摘要

被引文献

相似文献

每天都会产生大量的数字文本信息。有效地检索、管理和开发文本数据已成为当前的主要任务。在本文中,我们首先介绍了文本挖掘和概率主题模型潜狄利克雷分配。然后提出了两个实验——维基百科文章和用户推文主题建模。前者建立了一个文档主题模型,旨在从主题的角度解决文章的搜索、挖掘和推荐问题。后者建立了用户话题模型,对Twitter用户的兴趣进行了全面的研究和分析。实验过程包括数据收集、数据预处理和模型训练,并进行了完整的记录和评论。此外,本文的结论和应用可以为社会和商业研究提供有用的计算工具。
A Large number of digital text information is generated every day. Effectively searching, managing and exploring the text data has become a main task. In this paper, we first represent an introduction to text mining and a probabilistic topic model Latent Dirichlet allocation. Then two experiments are proposed - Wikipedia articles and users’ tweets topic modelling. The former one builds up a document topic model, aiming to a topic perspective solution on searching, exploring and recommending articles. The latter one sets up a user topic model, providing a full research and analysis over Twitter users’ interest. The experiment process including data collecting, data pre-processing and model training is fully documented and commented. Further more, the conclusion and application of this paper could be a useful computation tool for social and business research.