Recovering traceability links between code and documentation

Recovering traceability links between code and documentation
复制标题

DOI:
10.1109/tse.2002.1041053
复制
发表时间:
2002-10-01
影响因子:
7.4
通讯作者:
Merlo, E
Merlo, E
中科院分区:
计算机科学1区
文献类型:
--
作者:
Antoniol, G;Canfora, G;Merlo, E

文献摘要

被引文献

相似文献

软件系统文档几乎总是用自然语言和自由文本非正式地表达。例子包括需求说明、设计文档、手册页、系统开发日志、错误日志和相关的维护报告。我们提出了一种基于信息检索的方法来恢复源代码和自由文本文档之间的可追溯性链接。我们工作的前提是程序员对程序项使用有意义的名称,例如函数、变量、类型、类和方法。我们认为,程序员在编写代码时处理的应用程序领域知识通常由标识符的助记符捕获;因此,对这些助记符的分析有助于将高级概念与程序概念联系起来,反之亦然。我们在两个案例研究中同时应用概率和向量空间信息检索模型,将c++源代码跟踪到手册页,将Java代码跟踪到功能需求。我们比较了应用这两种模型的结果,讨论了优点和局限性,并描述了改进的方向。
Software system documentation is almost always expressed informally in natural language and free text. Examples include requirement specifications, design documents, manual pages, system development journals, error logs, and related maintenance reports. We propose a method based on information retrieval to recover traceability links between source code and free text documents. A premise of our work is that programmers use meaningful names for program items, such as functions, variables, types, classes, and methods. We believe that the application-domain knowledge that programmers process when writing the code is often captured by the mnemonics for identifiers; therefore, the analysis of these mnemonics can help to associate high-level concepts with program concepts and vice-versa. We apply both a probabilistic and a vector space information retrieval model in two case studies to trace C++ source code onto manual pages and Java code to functional requirements. We compare the results of applying the two models, discuss the benefits and limitations, and describe directions for improvements.