课题基金 / 基金详情

Statistical Natural Language Processing Methods for Computer Program Source Code

Statistical Natural Language Processing Methods for Computer Program Source Code
计算机程序源代码的统计自然语言处理方法
批准号:
EP/K024043/1
负责人:
Charles Sutton
金额:
$47.86万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --

项目摘要

项目成果

Charles Sutton的其他基金

相似基金

相关文献

中文摘要
翻译
复杂的软件系统涉及许多组件,并使用许多外部库。使用这些软件的程序员必须记住正确使用所有这些组件的协议,学习使用新组件的过程可能会很耗时,并且是bug的来源。我们相信有一个尚未开发的资源可以帮助解决这个问题。互联网上有数十亿行代码,其中大部分具有专业质量。这些代码中隐藏着大量关于良好编码实践的知识,例如,关于避免容易出错的构造或关于使用特定库的最佳协议。我们设想了一种新型的编程工具,它可以被称为数据驱动的开发工具,从一个大型的语料库的成熟的软件项目的编程知识的聚合,在开发环境中的演示文稿。正如当前一代IDE帮助开发人员管理代码一样,下一代IDE将帮助开发人员学习如何编写更好的代码。幸运的是,有一个研究领域已经开发出大量用于分析大量文本的复杂工具:即统计自然语言处理。该项目的长期战略目标是开发新的自然语言处理技术,旨在分析计算机程序源代码,以帮助程序员从他人的代码中学习编码技术。这里有一个很大的研究领域几乎完全没有被探索过。作为这个研究领域的第一步,在这个项目中,我们将专注于自动识别在不同软件项目中重复出现的短代码片段,我们称之为习惯用法。习惯用法的一个例子是Java中迭代数组的典型构造。虽然它们在源代码中无处不在,但据我们所知,这种形式的习惯用法还没有被系统地研究过,我们也不知道有任何自动识别习惯用法的技术。该项目的主要目标是开发新的统计NLP方法,目标是从源代码文本语料库中自动识别习语。我们称这个研究问题为习语挖掘,据我们所知,这是一个新的研究问题。这是一个跨学科的项目,借鉴了统计NLP,机器学习和软件工程。该项目的研究工作主要是统计NLP和机器学习,并将涉及开发新的统计方法来发现编程语言文本中的习惯用法。
英文摘要
Complex software systems involve many components and make use of many external libraries. Programmers who work on such software must remember the protocols for using all of those components correctly, and the process of learning to use a new component can be time consuming and a source of bugs.We believe that there is a major untapped resource that can help address this problem. Billions of lines of code are readily available on the Internet, much of which are of professional quality. Hidden within this code is a large amount of knowledge about good coding practices, for example, about avoiding error-prone constructs or about the best protocol for using a particular library. We envision a new type of programming tool, which could be called data-driven development tools, that aggregate knowledge about programming from a large corpus of mature software projects, for presentation within the development environment. Just as the current generation of IDEs helps developers to manage their code, the next generation of IDEs will help developers to learn how to write better code.Fortunately, there is a research field that has already developed a large body of sophisticated tools for analyzing large amounts of text: namely, statistical natural language processing. The long-term strategic goal of this project is to develop new natural language processing techniques aimed at analyzing computer program source code, in order to help programmers learn coding techniques from the code of others. There is a large area for research here that has been almost completely unexplored.As a first step in this research area, in this project we will focus on automatically identifying short code fragments, which we call idioms, that occur repeatedly across different software projects. An example of an idiom is the typical construct for iterating over an array in Java. Although they are ubiquitous in source code, idioms of this form have not to our knowledge been systematically studied, and we are unaware of any techniques for automatically identifying idioms. The main objective of this project is to develop new statistical NLP methods with the goal of automatically identifying idioms from a corpus of source code text. We call this research problem idiom mining, and it is to our knowledge a new research problem.This is an interdisciplinary project that draws from statistical NLP, machine learning, and software engineering. The research work of this project is primarily in statistical NLP and machine learning, and will involve developing new statistical methods for finding idioms in programming language text.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3212695
发表时间: 2018-09-01
期刊: ACM COMPUTING SURVEYS
影响因子: 16.6
作者: [Allamanis, Miltiadis, Barr, Earl T., Sutton, Charles]
通讯作者: Sutton, Charles
Learning Natural Coding Conventions
学习自然编码约定
DOI: 10.48550/arxiv.1402.4182
发表时间: 2014
期刊: arXiv e-prints
影响因子: --
作者: [Allamanis Miltiadis]
通讯作者: Allamanis Miltiadis
DOI: --
发表时间: 2016-02
期刊: ArXiv
影响因子: --
作者: [Miltiadis Allamanis;Hao Peng;Charles Sutton]
通讯作者: Miltiadis Allamanis;Hao Peng;Charles Sutton
DOI: 10.1109/tse.2018.2832048
发表时间: 2018-07
期刊: IEEE Transactions on Software Engineering
影响因子: 7.4
作者: [Miltiadis Allamanis;Earl T. Barr;C. Bird;Premkumar T. Devanbu;Mark Marron;Charles Sutton]
通讯作者: Miltiadis Allamanis;Earl T. Barr;C. Bird;Premkumar T. Devanbu;Mark Marron;Charles Sutton
LUCID: Clearer Software by Integrating Natural Language Analysis into Software Engineering
  • 批准号:
    EP/P005314/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $39.08万
  • 财政年份:
    2017
  • 负责人:
    Charles Sutton
  • 依托单位:
Fast, Locally Adaptive Inference for Machine Learning in Graphical Models
  • 批准号:
    EP/J00104X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $11.94万
  • 财政年份:
    2011
  • 负责人:
    Charles Sutton
  • 依托单位:
国内基金
海外基金
Natural超对称中的希格斯物理与暗物质研究
  • 批准号:
    11775039
  • 项目类别:
    面上项目
  • 资助金额:
    52.0万元
  • 批准年份:
    2017
  • 负责人:
    郑思波
  • 依托单位:
Natural超对称在LHC上的现象学研究
  • 批准号:
    11405015
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2014
  • 负责人:
    郑思波
  • 依托单位: