Statistical Unigram Analysis for Source Code Repository
Statistical Unigram Analysis for Source Code Repository
复制标题
源代码存储库的统计一元分析
DOI:
10.1109/bigmm.2017.13
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Abdulrahman Alatawi
中科院分区:
文献类型:
--
作者:
Dianxiang Xu;O. Ariss;Yunkai Liu;Abdulrahman Alatawi
Unigram is a fundamental element of n-gram in natural language processing. However, unigrams collected from a natural language corpus are unsuitable for solving problems in the domain of computer programming languages. In this paper, we analyze the properties of unigrams collected from an ultra-large source code repository. Specifically, we have collected 1.01 billion unigrams from 0.7 million open source projects hosted at GitHub.com. By analyzing these unigrams, we have discovered statistical patterns regarding (1) how developers name variables, methods, and classes, and (2) how developers choose abbreviations. Our study describes a probabilistic model for solving a well-known problem in source code analysis: how to expand a given abbreviation to its original indented word. It shows that the unigrams collected from source code repositories are essential resources to solving the domain specific problems.