Automatically Identifying the Quality of Developer Chats for Post Hoc Use

Automatically Identifying the Quality of Developer Chats for Post Hoc Use
复制标题

DOI:
10.1145/3450503
复制
发表时间:
2021-07
期刊:
ACM Transactions on Software Engineering and Methodology (TOSEM)
影响因子:
--
通讯作者:
Preetha Chatterjee;Kostadin Damevski;Nicholas A. Kraft;L. Pollock
Preetha Chatterjee;Kostadin Damevski;Nicholas A. Kraft;L. Pollock
中科院分区:
其他
文献类型:
--
作者:
Preetha Chatterjee;Kostadin Damevski;Nicholas A. Kraft;L. Pollock

文献摘要

相似文献

软件工程师正在Q&A论坛上众包他们日常挑战的答案(例如,Stack Overflow)以及最近的公共聊天社区,如Slack,IRC和Gitter。许多与软件相关的聊天对话包含有价值的专业知识,这些知识对于挖掘以改进编程支持工具和没有参与原始聊天对话的读者都很有用。然而,大多数聊天平台和社区不包含内置的质量指标(例如,答案,投票数)。因此,很难识别包含用于挖掘或阅读的有用信息的对话,即,事后质量的谈话。在这篇文章中,我们研究了从公共聊天频道中自动检测事后质量的开发人员对话。我们首先描述了400个开发人员对话的分析,这些对话表明了事后质量的潜在特征,然后是一种基于机器学习的方法,用于自动识别事后质量的对话。我们对四个编程社区(python,clojure,elm和racket)中的2,000个带注释的Slack对话进行了评估,结果表明我们的方法可以实现0.82的精确度,0.90的召回率,0.86的F-测量值和0.57的MCC。据我们所知,这是第一个用于检测事后质量的开发人员对话的自动化技术。
Software engineers are crowdsourcing answers to their everyday challenges on Q&A forums (e.g., Stack Overflow) and more recently in public chat communities such as Slack, IRC, and Gitter. Many software-related chat conversations contain valuable expert knowledge that is useful for both mining to improve programming support tools and for readers who did not participate in the original chat conversations. However, most chat platforms and communities do not contain built-in quality indicators (e.g., accepted answers, vote counts). Therefore, it is difficult to identify conversations that contain useful information for mining or reading, i.e., conversations of post hoc quality. In this article, we investigate automatically detecting developer conversations of post hoc quality from public chat channels. We first describe an analysis of 400 developer conversations that indicate potential characteristics of post hoc quality, followed by a machine learning-based approach for automatically identifying conversations of post hoc quality. Our evaluation of 2,000 annotated Slack conversations in four programming communities (python, clojure, elm, and racket) indicates that our approach can achieve precision of 0.82, recall of 0.90, F-measure of 0.86, and MCC of 0.57. To our knowledge, this is the first automated technique for detecting developer conversations of post hoc quality.