Predicting Query Quality for Applications of Text Retrieval to Software Engineering Tasks

Predicting Query Quality for Applications of Text Retrieval to Software Engineering Tasks
复制标题

预测文本检索在软件工程任务中的应用的查询质量

DOI:
10.1145/3078841
复制
发表时间:
2017
期刊:
ACM Transactions on Software Engineering and Methodology (TOSEM)
影响因子:
--
通讯作者:
A. D. Lucia
A. D. Lucia
中科院分区:
--
文献类型:
--
作者:
Chris Mills;G. Bavota;S. Haiduc;Rocco Oliveto;Andrian Marcus;A. D. Lucia

文献摘要

被引文献

相似文献

内容:自2000年代中期以来,大量基于文本检索(TR)的推荐系统被提出来支持软件工程(SE)任务,如概念定位、可追溯性链接恢复、代码重用、影响分析等,其成功与否高度依赖于提交的查询,这些查询要么是由开发人员制定的,要么是从软件工件中自动提取的。目的:我们的目标是预测质量的查询提交到TR为基础的方法在SE。这可以为开发人员和软件系统的质量带来好处。例如,了解查询何时表述不当可以为开发人员节省分析不相关搜索结果的时间和挫折感。相反,他们可以专注于重新定义查询。此外,知道用作查询的工件是否导致不相关的搜索结果可以揭示查询工件本身中的潜在问题。方法:我们引入了一个自动查询质量预测的方法,软件工件检索适应自然语言启发的解决方案,他们对软件数据的使用。我们提出了两个应用程序和评估的方法的背景下,概念定位和可追溯性链接恢复,TR已被应用最经常在SE。对于概念位置,我们使用的方法来确定是否检索到的代码元素的列表可能包含代码相关的一个特定的更改请求或没有,在这种情况下,查询是很好的候选人重新制定。对于可追溯性链接恢复,查询表示软件工件。在这种情况下,我们使用查询质量预测方法来识别难以跟踪到其他工件的工件,因此可能具有基于TR的可跟踪性链接恢复的低内在质量。结果如下:对于概念定位,评估表明,我们的方法能够在82%的情况下正确预测查询的质量,平均而言,使用很少的训练数据。在可追溯性恢复的情况下,所提出的方法是能够检测到难以跟踪的文物在74%的情况下,平均。结论:我们对概念定位和可追溯性链接恢复应用程序的评估结果表明,我们的方法可以用来预测基于TR的方法通过评估文本查询的质量的结果。这可以节省精力和时间,并识别可能难以使用TR跟踪的软件工件。
Context: Since the mid-2000s, numerous recommendation systems based on text retrieval (TR) have been proposed to support software engineering (SE) tasks such as concept location, traceability link recovery, code reuse, impact analysis, and so on. The success of TR-based solutions highly depends on the query submitted, which is either formulated by the developer or automatically extracted from software artifacts. Aim: We aim at predicting the quality of queries submitted to TR-based approaches in SE. This can lead to benefits for developers and for the quality of software systems alike. For example, knowing when a query is poorly formulated can save developers the time and frustration of analyzing irrelevant search results. Instead, they could focus on reformulating the query. Also, knowing if an artifact used as a query leads to irrelevant search results may uncover underlying problems in the query artifact itself. Method: We introduce an automatic query quality prediction approach for software artifact retrieval by adapting NL-inspired solutions to their use on software data. We present two applications and evaluations of the approach in the context of concept location and traceability link recovery, where TR has been applied most often in SE. For concept location, we use the approach to determine if the list of retrieved code elements is likely to contain code relevant to a particular change request or not, in which case, the queries are good candidates for reformulation. For traceability link recovery, the queries represent software artifacts. In this case, we use the query quality prediction approach to identify artifacts that are hard to trace to other artifacts and may therefore have a low intrinsic quality for TR-based traceability link recovery. Results: For concept location, the evaluation shows that our approach is able to correctly predict the quality of queries in 82% of the cases, on average, using very little training data. In the case of traceability recovery, the proposed approach is able to detect hard to trace artifacts in 74% of the cases, on average. Conclusions: The results of our evaluation on applications for concept location and traceability link recovery indicate that our approach can be used to predict the results of a TR-based approach by assessing the quality of the text query. This can lead to saved effort and time, as well as the identification of software artifacts that may be difficult to trace using TR.