Evaluating the Coverage and Depth of Latent Dirichlet Allocation Topic Model in Comparison with Human Coding of Qualitative Data: The Case of Education Research

Evaluating the Coverage and Depth of Latent Dirichlet Allocation Topic Model in Comparison with Human Coding of Qualitative Data: The Case of Education Research
复制标题

DOI:
10.3390/make5020029
复制
发表时间:
2023-05
期刊:
Mach. Learn. Knowl. Extr.
影响因子:
--
通讯作者:
Gaurav Nanda;A. Jaiswal;Hugo Castellanos;Yuzhe Zhou;Alex Choi;Alejandra J. Magana
Gaurav Nanda;A. Jaiswal;Hugo Castellanos;Yuzhe Zhou;Alex Choi;Alejandra J. Magana
中科院分区:
其他
文献类型:
--
作者:
Gaurav Nanda;A. Jaiswal;Hugo Castellanos;Yuzhe Zhou;Alex Choi;Alejandra J. Magana

文献摘要

被引文献

相似文献

社会科学领域,如教育研究,已开始扩大使用基于计算机的研究方法,以补充传统的研究方法。自然语言处理技术,如主题建模,可以通过提供研究人员可以解释和细化的早期类别来支持定性数据分析。这项研究有助于这一机构的研究,并回答了以下研究问题:(RQ 1)什么是相对覆盖率的潜在狄利克雷分配(LDA)的主题模型和人类编码的广度的主题/主题提取的文本集合?(RQ2)使用LDA主题模型和人工编码方法识别的主题之间的相对深度或详细程度是多少?使用LDA主题建模和人类编码方法对学生反思的数据集进行了定性分析,并对结果进行了比较。研究结果表明,主题模型可以提供可靠的覆盖面和深度的主题,目前在一个文本集合可比的人类编码,但需要手动解释的主题。人类编码输出的广度和深度在很大程度上取决于编码人员的专业知识和集合的大小;这些因素在主题建模方法中得到了更好的处理。
Fields in the social sciences, such as education research, have started to expand the use of computer-based research methods to supplement traditional research approaches. Natural language processing techniques, such as topic modeling, may support qualitative data analysis by providing early categories that researchers may interpret and refine. This study contributes to this body of research and answers the following research questions: (RQ1) What is the relative coverage of the latent Dirichlet allocation (LDA) topic model and human coding in terms of the breadth of the topics/themes extracted from the text collection? (RQ2) What is the relative depth or level of detail among identified topics using LDA topic models and human coding approaches? A dataset of student reflections was qualitatively analyzed using LDA topic modeling and human coding approaches, and the results were compared. The findings suggest that topic models can provide reliable coverage and depth of themes present in a textual collection comparable to human coding but require manual interpretation of topics. The breadth and depth of human coding output is heavily dependent on the expertise of coders and the size of the collection; these factors are better handled in the topic modeling approach.