The Expansion of Source Code Abbreviations Using a Language Model
The Expansion of Source Code Abbreviations Using a Language Model
复制标题
使用语言模型扩展源代码缩写
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Jie Yan
中科院分区:
文献类型:
--
作者:
Abdulrahman Alatawi;Weifeng Xu;Jie Yan
Programmers often abbreviate identifiers names in source code to represent single words, i.e. unigrams, or phrases, i.e. multigrams. However, the difficulty to retrieve the original word(s) of an abbreviation during the maintenance phase makes the source code more problematic to comprehend. Incorrect abbreviations expansion may lead to introducing defects in the code. There are many approaches that that automatically expand abbreviations to their original words, unfortunately, they are based on predefined patterns and single-words dictionaries which cannot address abbreviations that are expandable to phrases. In this paper, we describe a bigram-based inference model which utilizes unigrams statistical properties as evidence to retrieve the original word automatically. We evaluated our approach on a set of 100 abbreviations randomly picked from eight open source projects and found that our approach correctly expands 78% of the set.