Inferring semantically related words from software context

Inferring semantically related words from software context
复制标题

DOI:
10.1109/msr.2012.6224276
复制
发表时间:
2012-06
期刊:
2012 9th IEEE Working Conference on Mining Software Repositories (MSR)
影响因子:
--
通讯作者:
Jinqiu Yang;Lin Tan
Jinqiu Yang;Lin Tan
中科院分区:
其他
文献类型:
--
作者:
Jinqiu Yang;Lin Tan

文献摘要

被引文献

相似文献

代码搜索是软件开发和程序理解的一个组成部分。代码搜索的困难在于无法猜测代码中使用的确切单词。因此,对于基于关键字的代码搜索来说,用语义相关的词来扩展查询是至关重要的,例如,同义词和缩写,以提高搜索效率。然而,依靠英语词典和WordNet等资源来获取软件中的语义相关词是有限的,因为许多在软件中语义相关的词在英语中并不语义相关。本文提出了一种简单而通用的技术,通过利用注释和代码中单词的上下文来自动推断软件中语义相关的单词。我们实现了合理的准确性,在七个大型和流行的代码库编写的C和Java。通过对现有技术的进一步评价表明,该方法具有较高的查准率和查全率。
Code search is an integral part of software development and program comprehension. The difficulty of code search lies in the inability to guess the exact words used in the code. Therefore, it is crucial for keyword-based code search to expand queries with semantically related words, e.g., synonyms and abbreviations, to increase the search effectiveness. However, it is limited to rely on resources such as English dictionaries and WordNet to obtain semantically related words in software, because many words that are semantically related in software are not semantically related in English. This paper proposes a simple and general technique to automatically infer semantically related words in software by leveraging the context of words in comments and code. We achieve a reasonable accuracy in seven large and popular code bases written in C and Java. Our further evaluation against the state of art shows that our technique can achieve a higher precision and recall.