The Expansion of Source Code Abbreviations Using a Language Model

The Expansion of Source Code Abbreviations Using a Language Model
复制标题

使用语言模型扩展源代码缩写

DOI:
--
复制
发表时间:
2018
期刊:
Annual International Computer Software and Applications Conference
影响因子:
--
通讯作者:
Jie Yan
Jie Yan
中科院分区:
--
文献类型:
--
作者:
Abdulrahman Alatawi;Weifeng Xu;Jie Yan

文献摘要

被引文献

相似文献

程序员经常在源代码中使用标识符名称来表示单个单词(即一元语法)或短语(即多元语法)。然而,在维护阶段检索缩写的原始单词的困难使得源代码更难以理解。不正确的缩写扩展可能会导致在代码中引入缺陷。有许多方法可以自动将缩写扩展到其原始单词,不幸的是,它们基于预定义的模式和无法处理可扩展为短语的缩写的单字词典。在本文中,我们描述了一个基于bigram的推理模型,利用unigram的统计特性作为证据,自动检索原始词。我们对从8个开源项目中随机挑选的100个缩写进行了评估,发现我们的方法正确地扩展了78%的缩写。
Programmers often abbreviate identifiers names in source code to represent single words, i.e. unigrams, or phrases, i.e. multigrams. However, the difficulty to retrieve the original word(s) of an abbreviation during the maintenance phase makes the source code more problematic to comprehend. Incorrect abbreviations expansion may lead to introducing defects in the code. There are many approaches that that automatically expand abbreviations to their original words, unfortunately, they are based on predefined patterns and single-words dictionaries which cannot address abbreviations that are expandable to phrases. In this paper, we describe a bigram-based inference model which utilizes unigrams statistical properties as evidence to retrieve the original word automatically. We evaluated our approach on a set of 100 abbreviations randomly picked from eight open source projects and found that our approach correctly expands 78% of the set.