Detection of Duplicate Defect Reports Using Natural Language Processing

Detection of Duplicate Defect Reports Using Natural Language Processing
复制标题

DOI:
10.1109/icse.2007.32
复制
发表时间:
2007-05
期刊:
29th International Conference on Software Engineering (ICSE'07)
影响因子:
--
通讯作者:
P. Runeson;Magnus Alexandersson;Oskar Nyholm
P. Runeson;Magnus Alexandersson;Oskar Nyholm
中科院分区:
其他
文献类型:
--
作者:
P. Runeson;Magnus Alexandersson;Oskar Nyholm

文献摘要

被引文献

相似文献

缺陷报告是由软件工程中各种测试和开发活动产生的。有时,提交了两个报告相同问题的报告,从而导致重复报告。这些报告主要是用结构化的自然语言编写的,因此,很难将两个报告与形式方法进行比较。为了识别重复项,我们使用自然语言处理(NLP)技术进行研究以支持识别。在案例研究中开发和评估了原型工具,分析了索尼爱立信移动通信的缺陷报告。评估表明,使用NLP技术可能可以找到约2/3的重复项。这些技术的不同变体仅提供较小的结果差异,表明技术的技术稳定。用户测试表明,对该技术的总体态度是积极的,并且具有增长潜力。
Defect reports are generated from various testing and development activities in software engineering. Sometimes two reports are submitted that describe the same problem, leading to duplicate reports. These reports are mostly written in structured natural language, and as such, it is hard to compare two reports for similarity with formal methods. In order to identify duplicates, we investigate using natural language processing (NLP) techniques to support the identification. A prototype tool is developed and evaluated in a case study analyzing defect reports at Sony Ericsson mobile communications. The evaluation shows that about 2/3 of the duplicates can possibly be found using the NLP techniques. Different variants of the techniques provide only minor result differences, indicating a robust technology. User testing shows that the overall attitude towards the technique is positive and that it has a growth potential.