Statistical Unigram Analysis for Source Code Repository

Statistical Unigram Analysis for Source Code Repository
复制标题

源代码存储库的统计一元分析

DOI:
10.1109/bigmm.2017.13
复制
发表时间:
2017
期刊:
2017 IEEE Third International Conference on Multimedia Big Data (BigMM)
影响因子:
--
通讯作者:
Abdulrahman Alatawi
Abdulrahman Alatawi
中科院分区:
--
文献类型:
--
作者:
Dianxiang Xu;O. Ariss;Yunkai Liu;Abdulrahman Alatawi

文献摘要

被引文献

相似文献

Unigram是自然语言处理中n-gram的基本元素。然而,从自然语言语料库中收集的一元语法不适合用于解决计算机编程语言领域中的问题。在本文中,我们分析了从一个超大型源代码库中收集的unigrams的属性。具体来说,我们已经从GitHub.com上托管的70万个开源项目中收集了10.1亿个unigram。通过分析这些unigrams,我们发现了关于(1)开发人员如何命名变量,方法和类,以及(2)开发人员如何选择缩写的统计模式。我们的研究描述了一个概率模型,用于解决源代码分析中的一个众所周知的问题:如何将给定的缩写扩展到其原始的缩进单词。它表明,从源代码存储库中收集的unigram是解决特定领域问题的必要资源。
Unigram is a fundamental element of n-gram in natural language processing. However, unigrams collected from a natural language corpus are unsuitable for solving problems in the domain of computer programming languages. In this paper, we analyze the properties of unigrams collected from an ultra-large source code repository. Specifically, we have collected 1.01 billion unigrams from 0.7 million open source projects hosted at GitHub.com. By analyzing these unigrams, we have discovered statistical patterns regarding (1) how developers name variables, methods, and classes, and (2) how developers choose abbreviations. Our study describes a probabilistic model for solving a well-known problem in source code analysis: how to expand a given abbreviation to its original indented word. It shows that the unigrams collected from source code repositories are essential resources to solving the domain specific problems.