Empirical Principles and an Industrial Case Study in Retrieving Equivalent Requirements via Natural Language Processing Techniques

Empirical Principles and an Industrial Case Study in Retrieving Equivalent Requirements via Natural Language Processing Techniques
复制标题

DOI:
10.1109/tse.2011.122
复制
发表时间:
2013
影响因子:
7.4
通讯作者:
D. Falessi;G. Cantone;G. Canfora
D. Falessi;G. Cantone;G. Canfora
中科院分区:
计算机科学1区
文献类型:
--
作者:
D. Falessi;G. Cantone;G. Canfora

文献摘要

被引文献

相似文献

虽然在软件工程中非常重要,但是链接相同类型(克隆检测)或不同类型(可跟踪性恢复)的工件是非常乏味、容易出错和工作量大的。过去的研究主要集中在支持分析师使用基于自然语言处理(NLP)的技术来识别候选链接。由于存在许多NLP技术,并且它们的性能根据上下文而变化,因此定义和使用可靠的评估程序至关重要。本文的目的是提出一套七项原则,用于评估NLP技术在识别等效需求方面的性能。在本文中,我们推测,并验证,NLP技术执行一个给定的数据集,根据能力和正确识别等价要求的几率。例如,当识别等价需求的几率非常高时,那么就有理由期望NLP技术会产生良好的性能。我们的关键思想是测量使用中的特定数据集的随机因子,然后相应地调整观察到的性能。为了支持应用的原则,我们报告其实际应用的案例研究,评估了大量的自然语言处理技术的性能,以确定在国防和航空航天领域的意大利公司的背景下的等效要求。当前的应用环境是评估NLP技术以识别等效需求。然而,大多数提出的原理似乎适用于评估旨在支持二元决策的任何估计技术(例如,等价/不等价),估计值在[0,1]范围内(例如,由NLP提供的相似性),当数据集被用作基准时(即,测试床),与估计器的类型无关(即,要求文本)和估计方法(例如,NLP)。
Though very important in software engineering, linking artifacts of the same type (clone detection) or different types (traceability recovery) is extremely tedious, error-prone, and effort-intensive. Past research focused on supporting analysts with techniques based on Natural Language Processing (NLP) to identify candidate links. Because many NLP techniques exist and their performance varies according to context, it is crucial to define and use reliable evaluation procedures. The aim of this paper is to propose a set of seven principles for evaluating the performance of NLP techniques in identifying equivalent requirements. In this paper, we conjecture, and verify, that NLP techniques perform on a given dataset according to both ability and the odds of identifying equivalent requirements correctly. For instance, when the odds of identifying equivalent requirements are very high, then it is reasonable to expect that NLP techniques will result in good performance. Our key idea is to measure this random factor of the specific dataset(s) in use and then adjust the observed performance accordingly. To support the application of the principles we report their practical application to a case study that evaluates the performance of a large number of NLP techniques for identifying equivalent requirements in the context of an Italian company in the defense and aerospace domain. The current application context is the evaluation of NLP techniques to identify equivalent requirements. However, most of the proposed principles seem applicable to evaluating any estimation technique aimed at supporting a binary decision (e.g., equivalent/nonequivalent), with the estimate in the range [0,1] (e.g., the similarity provided by the NLP), when the dataset(s) is used as a benchmark (i.e., testbed), independently of the type of estimator (i.e., requirements text) and of the estimation method (e.g., NLP).