A Topic Modeling Comparison Between LDA, NMF, Top2Vec, and BERTopic to Demystify Twitter Posts.

A Topic Modeling Comparison Between LDA, NMF, Top2Vec, and BERTopic to Demystify Twitter Posts.
复制标题

DOI:
10.3389/fsoc.2022.886498
复制
发表时间:
2022
影响因子:
2.5
通讯作者:
Yu, Joanne
Yu, Joanne
中科院分区:
其他
文献类型:
--
作者:
Egger, Roman;Yu, Joanne

文献摘要

被引文献

相似文献

社交媒体数据的丰富性为社会科学研究开辟了一条新的途径,以深入了解人类的行为和经验。特别是,依赖于主题模型的新兴数据驱动方法为解释社会现象提供了全新的视角。然而,社交媒体内容的简短,文本密集和非结构化的性质往往导致数据收集和分析的方法挑战。为了弥合计算科学和实证社会研究的发展领域,本研究旨在评估四个主题建模技术的性能,即潜在狄利克雷分配(LDA),非负矩阵分解(NMF),Top2Vec,和BERTopic。鉴于人际关系和数字媒体之间的相互作用,本研究以Twitter帖子为参考点,并评估了不同算法在社会科学背景下的优缺点。基于分析过程中的某些细节和质量问题,本研究揭示了使用BERTopic和NMF分析Twitter数据的有效性。
The richness of social media data has opened a new avenue for social science research to gain insights into human behaviors and experiences. In particular, emerging data-driven approaches relying on topic models provide entirely new perspectives on interpreting social phenomena. However, the short, text-heavy, and unstructured nature of social media content often leads to methodological challenges in both data collection and analysis. In order to bridge the developing field of computational science and empirical social research, this study aims to evaluate the performance of four topic modeling techniques; namely latent Dirichlet allocation (LDA), non-negative matrix factorization (NMF), Top2Vec, and BERTopic. In view of the interplay between human relations and digital media, this research takes Twitter posts as the reference point and assesses the performance of different algorithms concerning their strengths and weaknesses in a social science context. Based on certain details during the analytical procedures and on quality issues, this research sheds light on the efficacy of using BERTopic and NMF to analyze Twitter data.