An effective hot topic detection method for microblog on spark

An effective hot topic detection method for microblog on spark
复制标题

一种有效的Spark微博热点话题检测方法

DOI:
10.1016/j.asoc.2017.08.053
复制
发表时间:
2018-09-01
影响因子:
8.7
通讯作者:
Li, Keqin
Li, Keqin
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ai, Wei;Li, Kenli;Li, Keqin

文献摘要

被引文献

相似文献

随着大数据时代的出现,从海量的数字化文本材料中快速、准确地获取有价值的热点话题的方法备受关注。在这项工作中,我们专注于大数据环境下微博的主题检测。与现有方法不同,我们以分布式方式解决这个问题。具体来说,我们提出了一种称为并行两相麦克风-麦克热点主题检测(TMHTD)的非迭代算法,并在 Apache Spark 环境中实现它。所提出的TMHTD方法包括两个阶段,即微聚类阶段和宏聚类阶段。为了提高热点主题检测的准确性,提出了三种优化方法以及TMHTD。为了处理大型数据库,我们特意设计了一组MapReduce作业,以高度可扩展的方式具体完成热点主题检测。我们将TMHTD算法与一般的单遍算法和潜在狄利克雷分配(LDA)算法进行比较。我们的实验是在从新浪微博 API 收集的现实数据集上进行的。大量的实验结果表明,TMHTD算法的精度和性能较之前的方法有显着的提高。更具体地,TMHTD算法的F-measure值分别比一般单通道算法和LDA算法提高了6%和8%。 TMHTD算法的运行时间分别比一般单遍算法和LDA算法优越7倍和2倍。 (C) 2017 年由 Elsevier B.V. 出版
With the emergence of the big data age, methods for quickly and accurately obtaining valuable hot topics from the vast amount of digitized textual material have attracted much attention. In this work, we focus on topic detection in microblogs in the big data environment. Different from existing approaches, we solve this problem in a distributed way. Specifically, we propose a non-iterative algorithm called parallel two-phase mic-mac hot topic detection (TMHTD), and implement it in the Apache Spark environment. The proposed TMHTD method includes two phases, i.e., the micro-clustering phase and the macro-clustering phase. To improve the accuracy of hot topic detection, three optimization methods, along with TMHTD, are proposed. To handle large databases, we deliberately design a group of MapReduce jobs to concretely accomplish hot topic detection in a highly scalable way. We compare the TMHTD algorithm with the general single-pass algorithm and the Latent Dirichlet Allocation (LDA) algorithm. Our experiments are carried out on real-life data sets gathered from the Sina Weibo API. Extensive experimental results indicate that the accuracy and performance of the TMHTD algorithm are significant improvements over previous methods. More specifically, the F-measure value of the TMHTD algorithm shows a 6% and 8% improvement over the general single-pass algorithm and the LDA algorithm, respectively. The run time of the TMHTD algorithm is 7 times and twice as superior to the general single-pass algorithm and the LDA algorithm, respectively. (C) 2017 Published by Elsevier B.V.