Statistical Natural Language Processing Methods for Computer Program Source Code
Statistical Natural Language Processing Methods for Computer Program Source Code
批准号:
EP/K024043/1
负责人:
Charles Sutton
金额:
$47.86万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --
中文摘要
复杂的软件系统涉及许多组件,并使用许多外部库。从事这类软件工作的程序员必须记住正确使用所有这些组件的协议,学习使用新组件的过程可能很耗时,而且会产生错误。我们相信,有一种主要的未开发资源可以帮助解决这个问题。数十亿行代码在互联网上随处可见,其中大部分是专业质量的。在这些代码中隐藏了大量关于良好编码实践的知识,例如,关于避免容易出错的构造或关于使用特定库的最佳协议。我们设想了一种新型的编程工具,它可以被称为数据驱动开发工具,它可以从大量成熟的软件项目中聚合关于编程的知识,以便在开发环境中呈现。正如当前一代IDE帮助开发人员管理他们的代码一样,下一代IDE将帮助开发人员学习如何编写更好的代码。幸运的是,有一个研究领域已经开发了大量用于分析大量文本的复杂工具:即统计自然语言处理。该项目的长期战略目标是开发新的自然语言处理技术,旨在分析计算机程序源代码,以帮助程序员从其他人的代码中学习编码技术。这里有一个几乎完全未开发的研究领域。作为这个研究领域的第一步,在这个项目中,我们将专注于自动识别短代码片段,我们称之为习惯用法,这些代码片段在不同的软件项目中重复出现。Java中用于迭代数组的典型结构就是一种习惯用法。尽管它们在源代码中无处不在,但据我们所知,这种形式的习语还没有被系统地研究过,我们也不知道任何自动识别习语的技术。该项目的主要目标是开发新的统计自然语言处理方法,目的是从源代码文本语料库中自动识别成语。我们将这个研究问题称为习语挖掘,据我们所知,它是一个新的研究问题。这是一个结合了统计自然语言处理、机器学习和软件工程的跨学科项目。该项目的研究工作主要是统计自然语言处理和机器学习,并将涉及开发新的统计方法,以寻找编程语言文本中的习语。
英文摘要
Complex software systems involve many components and make use of many external libraries. Programmers who work on such software must remember the protocols for using all of those components correctly, and the process of learning to use a new component can be time consuming and a source of bugs.We believe that there is a major untapped resource that can help address this problem. Billions of lines of code are readily available on the Internet, much of which are of professional quality. Hidden within this code is a large amount of knowledge about good coding practices, for example, about avoiding error-prone constructs or about the best protocol for using a particular library. We envision a new type of programming tool, which could be called data-driven development tools, that aggregate knowledge about programming from a large corpus of mature software projects, for presentation within the development environment. Just as the current generation of IDEs helps developers to manage their code, the next generation of IDEs will help developers to learn how to write better code.Fortunately, there is a research field that has already developed a large body of sophisticated tools for analyzing large amounts of text: namely, statistical natural language processing. The long-term strategic goal of this project is to develop new natural language processing techniques aimed at analyzing computer program source code, in order to help programmers learn coding techniques from the code of others. There is a large area for research here that has been almost completely unexplored.As a first step in this research area, in this project we will focus on automatically identifying short code fragments, which we call idioms, that occur repeatedly across different software projects. An example of an idiom is the typical construct for iterating over an array in Java. Although they are ubiquitous in source code, idioms of this form have not to our knowledge been systematically studied, and we are unaware of any techniques for automatically identifying idioms. The main objective of this project is to develop new statistical NLP methods with the goal of automatically identifying idioms from a corpus of source code text. We call this research problem idiom mining, and it is to our knowledge a new research problem.This is an interdisciplinary project that draws from statistical NLP, machine learning, and software engineering. The research work of this project is primarily in statistical NLP and machine learning, and will involve developing new statistical methods for finding idioms in programming language text.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3212695
发表时间:
2018-09-01
期刊:
ACM COMPUTING SURVEYS
影响因子:
16.6
作者:
[Allamanis, Miltiadis, Barr, Earl T., Sutton, Charles]
通讯作者:
Sutton, Charles
Learning Natural Coding Conventions
学习自然编码约定
DOI:
10.48550/arxiv.1402.4182
发表时间:
2014
期刊:
arXiv e-prints
影响因子:
--
作者:
[Allamanis Miltiadis]
通讯作者:
Allamanis Miltiadis
DOI:
--
发表时间:
2016-02
期刊:
ArXiv
影响因子:
--
作者:
[Miltiadis Allamanis;Hao Peng;Charles Sutton]
通讯作者:
Miltiadis Allamanis;Hao Peng;Charles Sutton
DOI:
10.1109/tse.2018.2832048
发表时间:
2018-07
期刊:
IEEE Transactions on Software Engineering
影响因子:
7.4
作者:
[Miltiadis Allamanis;Earl T. Barr;C. Bird;Premkumar T. Devanbu;Mark Marron;Charles Sutton]
通讯作者:
Miltiadis Allamanis;Earl T. Barr;C. Bird;Premkumar T. Devanbu;Mark Marron;Charles Sutton
DOI:
10.48550/arxiv.1611.01423
发表时间:
2016
期刊:
影响因子:
--
作者:
[Allamanis M]
通讯作者:
Allamanis M
LUCID: Clearer Software by Integrating Natural Language Analysis into Software Engineering
-
批准号:EP/P005314/1
-
项目类别:Research Grant
-
资助金额:$39.08万
-
财政年份:2017
-
负责人:Charles Sutton
-
依托单位:
Fast, Locally Adaptive Inference for Machine Learning in Graphical Models
-
批准号:EP/J00104X/1
-
项目类别:Research Grant
-
资助金额:$11.94万
-
财政年份:2011
-
负责人:Charles Sutton
-
依托单位:
国内基金
海外基金
Natural超对称中的希格斯物理与暗物质研究
-
批准号:11775039
-
项目类别:面上项目
-
资助金额:52.0万元
-
批准年份:2017
-
负责人:郑思波
-
依托单位:
Natural超对称在LHC上的现象学研究
-
批准号:11405015
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2014
-
负责人:郑思波
-
依托单位: