Topic Extraction and Classification for Questions Posted in Community-Based Question Answering Services

Topic Extraction and Classification for Questions Posted in Community-Based Question Answering Services
复制标题

DOI:
10.1109/csci49370.2019.00253
复制
发表时间:
2019-12
期刊:
2019 International Conference on Computational Science and Computational Intelligence (CSCI)
影响因子:
--
通讯作者:
Qing Ma;M. Murata
Qing Ma;M. Murata
中科院分区:
其他
文献类型:
--
作者:
Qing Ma;M. Murata

文献摘要

相似文献

本文介绍了使用主题模型和混合模型,同时介绍了基于社区的问答服务(CQA)或问答网站中发布的问题(CQA)或问答网站中发布的问题的方法。大规模实验对两种数据(一个称为类别数据,另一个称为亚型数据)显示了我们方法的有效性。纯度和正确的速率表明,该主题模型优于群集方法,混合模型优于所讨论的主题模型,并且采用术语频率内文档频率对于子类型数据有效。用提取的关键字进行的手动评估显示了主题模型在主题提取中的有效性。
This paper presents methods of simultaneously performing topic/keyword extraction and unsupervised classification for questions posted in community-based question answering services (CQA) or Q&A websites, using topic models and hybrid models. Large-scale experiments on two kinds of data, one called category data and the other called subtyping data, show the effectiveness of our methods. The purity and correct rate show that the topic models outperform clustering methods, hybrid models outperform topic models in question classification, and the adoption of term frequency-inverse document frequency is effective for the subtyping data. Manual evaluations with the extracted keywords show the effectiveness of the topic models in topic extraction.