A Survey of Machine Learning for Big Code and Naturalness

A Survey of Machine Learning for Big Code and Naturalness
复制标题

DOI:
10.1145/3212695
复制
发表时间:
2018-09-01
影响因子:
16.6
通讯作者:
Sutton, Charles
Sutton, Charles
中科院分区:
计算机科学1区
文献类型:
--
作者:
Allamanis, Miltiadis;Barr, Earl T.;Sutton, Charles

文献摘要

被引文献

相似文献

机器学习、编程语言和软件工程交叉领域的研究最近在提出可学习的源代码概率模型方面迈出了重要的一步,这些模型利用了大量的代码模式。在本文中,我们对这项工作进行了综述。我们将编程语言与自然语言进行对比,并讨论这些相似性和差异如何驱动概率模型的设计。我们根据每个模型的基本设计原则提出了一个分类法,并用它来浏览文献。然后,我们回顾了研究人员如何将这些模型应用于应用领域,并讨论了跨领域和特定应用的挑战和机遇。
Research at the intersection of machine learning, programming languages, and software engineering has recently taken important steps in proposing learnable probabilistic models of source code that exploit the abundance of patterns of code. In this article, we survey this work. We contrast programming languages against natural languages and discuss how these similarities and differences drive the design of probabilistic models. We present a taxonomy based on the underlying design principles of eachmodel and use it to navigate the literature. Then, we review how researchers have adapted these models to application areas and discuss cross-cutting and application-specific challenges and opportunities.