课题基金 / 基金详情

CRII: SHF: Towards the Construction of a Model for Natural Language and Source Code

CRII: SHF: Towards the Construction of a Model for Natural Language and Source Code
CRII:SHF:构建自然语言和源代码模型
批准号:
1850412
负责人:
Christian Newman
金额:
$17.45万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-05-01 至 2022-04-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
源代码使用人类语言(如英语)和编程语言的组合编写。开发人员结合使用人类语言和编程语言的规则来理解代码。试图理解代码的行为被称为程序理解;它是开发人员在编码时可能进行的所有其他与编程相关的活动之前的活动。例如,在修复错误之前,开发人员需要了解出现错误的代码;要添加新的软件功能,开发人员必须了解支持新功能的代码。如果一段代码是高度可理解的,那么开发人员将更容易地维护、调试和添加它。为了支持理解,研究必须尝试对人类语言如何描述程序行为进行正式建模。有了这样的模型,源代码可以通过自动改进或生成最好地描述它的人类语言来优化到最大限度地易于理解。本项目旨在通过将来自自然语言词性的信息与程序行为模型相结合来构建这样一个模型,以帮助、提高和测量理解。该项目旨在对人类语言描述源代码行为的方式进行正式建模。这将通过将基于静态分析的标识符类型分类与自然语言技术和标识符定义-使用链相结合来实现。这三个活动的组合允许模型测量1)类型如何约束标识符的行为,2)在英语中,标识符中的单词与什么角色相关,以及3)在哪个函数中使用该标识符。这些将允许模型理解标识符的英语如何与用法(函数调用)和行为约束(类型约束)相关。这个模型的目标是正式衡量人类语言被用来描述源代码行为的方式,这样它就可以用来训练机器做同样的事情。完成的模型将增加对开发人员如何通过人类语言表达程序行为的当前理解,并允许对该表达进行可测量的优化,以提高可理解性。此外,该模型将通过允许他们更多地了解基础源代码结构和规则如何影响使用人类语言描述程序行为的方式来改进现代程序理解技术。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Source code is written using a combination of human languages, such as English, and programming languages. Developers use a combination of the rules for human languages and programming languages to understand code. The act of trying to understand code is referred to as program comprehension; it is an activity that precedes all other programming-related activities a developer might undertake when coding. For example, before fixing a bug, a developer needs to understand the code where the bug is present; to add a new software feature, a developer must understand the code which will support the new feature. If a piece of code is highly comprehensible, then developers will have an easier time maintaining, debugging, and adding to it. To support comprehension, research must attempt to formally model how human language describes program behavior. With such a model, source code could be optimized to be maximally understandable by automatically improving, or generating, human language to best describe it. This project aims to build such a model by combining information from natural language part of speech with a model of program behavior to assist, improve and measure comprehension. This project aims to formally model how human language describes source code behavior. This will be achieved by combining a static-analysis-based taxonomy of identifier type categorizations with natural language techniques and identifier definition-use chains. The combination of these three activities allow the model to measure 1) how the type constrains the behavior of an identifier, 2) what role, in English, the words in an identifier correlate to, and 3) what function calls the identifier is used in. These will allow the model to understand how the English of an identifier relates to the usage (function calls) and behavior constraints (type constraints). The goal of this model is to formally measure the way human languages are used to describe source code behavior such that it could be used to train a machine to do the same. The completed model will increase the current understanding of how developers express program behavior through human languages and allow for this expression to be measurably optimized for increased comprehensibility. Additionally, the model will improve modern program comprehension techniques by allowing them to be more aware of how the underlying source code structure and rules influence the way human languages are used to describe program behavior.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
An Open Dataset of Abbreviations and Expansions
缩写和扩展的开放数据集
DOI: 10.1109/icsme.2019.00041
发表时间: 2019
期刊: International Conference on Software Maintenance and Evolution
影响因子: --
作者: [Newman, Christian, Decker, Michael John, AlSuhaibani, Reem S, Peruma, Anthony, Kaushik, Dishant, Hill, Emily]
通讯作者: Hill, Emily
IDEAL: An Open-Source Identifier Name Appraisal Tool
IDEAL:开源标识符名称评估工具
DOI: 10.1109/icsme52107.2021.00064
发表时间: 2021
期刊: IEEE International Conference on Software Maintenance and Evolution (ICSME
影响因子: --
作者: [Peruma, Anthony, Arnaoudova, Venera, Newman, Christian D.]
通讯作者: Newman, Christian D.
DOI: 10.1016/j.jss.2020.110740
发表时间: 2020-07
期刊: J. Syst. Softw.
影响因子: --
作者: [Christian D. Newman;Reem S. Alsuhaibani;M. J. Decker;Anthony S Peruma;D. Kaushik;Mohamed Wiem Mkaouer]
通讯作者: Christian D. Newman;Reem S. Alsuhaibani;M. J. Decker;Anthony S Peruma;D. Kaushik;Mohamed Wiem Mkaouer
Modeling the Relationship Between Identifier Name and Behavior
对标识符名称和行为之间的关系进行建模
DOI: 10.1109/icsme.2019.00062
发表时间: 2019
期刊: International Conference on Software Maintenance
影响因子: --
作者: [Newman, Christian D., Preuma, Anthony, AlSuhaibani, Reem]
通讯作者: AlSuhaibani, Reem
共 9 条
    国内基金
    海外基金
    天然超短抗菌肽Temporin-SHf衍生多肽的构效分析与抗菌机制研究
    衔接蛋白SHF负向调控胶质母细胞瘤中EGFR/EGFRvIII再循环和稳定性的功能及机制研究
    • 批准号:
      82302939
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2023
    • 负责人:
      汪京京
    • 依托单位:
    EGFR/GRβ/Shf调控环路在胶质瘤中的作用机制研究
    • 批准号:
      81572468
    • 项目类别:
      面上项目
    • 资助金额:
      60.0万元
    • 批准年份:
      2015
    • 负责人:
      邹健
    • 依托单位: