Understanding and Classifying the Quality of Technical Forum Questions

Understanding and Classifying the Quality of Technical Forum Questions
复制标题

了解技术论坛问题的质量并对其进行分类

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Quality Software
影响因子:
--
通讯作者:
Michele Lanza
Michele Lanza
中科院分区:
--
文献类型:
--
作者:
Luca Ponzanelli;Andrea Mocci;Alberto Bacchelli;Michele Lanza

文献摘要

被引文献

相似文献

技术问答(Q&A)服务已经成为开发人员的宝贵资源。技术问答网站的一个突出例子是StackOverflow (SO),它依赖于一个不断增长的社区,有超过200万的用户通过提问和提供答案积极地做出贡献。为了保持这一资源的价值,在每天提出的6000多个问题中,质量差的问题必须被过滤掉。目前,在SO中,低质量的问题是由选定的用户手动识别和审查的,这需要花费大量的时间和精力。自动化该过程将节省时间并卸载审查队列,提高SO作为开发人员资源的效率。我们提出了一种根据问题质量自动分类问题的方法。我们提出了一项实证研究,研究了如何通过考虑帖子内容(例如,从简单的文本特征到更复杂的可读性指标)和社区相关方面(例如,用户在社区中的受欢迎程度)的特征来建模和预测问题的质量。我们的研究结果表明,确实存在至少部分自动化昂贵的SO审查过程的可能性。
Technical questions and answers (Q&A) services have become a valuable resource for developers. A prominent example of technical Q&A website is StackOverflow (SO), which relies on a growing community of more than two millions of users who actively contribute by asking questions and providing answers. To maintain the value of this resource, poor quality questions - among the more than 6,000 asked daily - have to be filtered out. Currently, poor quality questions are manually identified and reviewed by selected users in SO, this costs considerable time and effort. Automating the process would save time and unload the review queue, improving the efficiency of SO as a resource for developers. We present an approach to automate the classification of questions according to their quality. We present an empirical study that investigates how to model and predict the quality of a question by considering as features both the contents of a post (e.g., from simple textual features to more complex readability metrics) and community-related aspects (e.g., popularity of a user in the community). Our findings show that there is indeed the possibility of at least a partial automation of the costly SO review process.