Deep Generation of Coq Lemma Names Using Elaborated Terms

Deep Generation of Coq Lemma Names Using Elaborated Terms
复制标题

DOI:
10.1007/978-3-030-51054-1_6
复制
发表时间:
2020-06-06
期刊:
Automated Reasoning
影响因子:
--
通讯作者:
Gligoric M
Gligoric M
中科院分区:
其他
文献类型:
--
作者:
Nie P;Palmskog K;Li JJ;Gligoric M

文献摘要

参考文献

相似文献

命名、间距和其他基本风格属性的编码约定对于开发人员有效地理解、审查和修改大型软件项目中的源代码是必要的。基于证明助手(如Coq)的验证项目中的一致约定随着项目规模和范围的增长而变得越来越重要。虽然约定可以以高成本手动记录和执行,但新兴的方法通过在大型代码语料库上应用统计语言模型来自动学习和建议Java类语言中的惯用名称。然而,由于其强大的语言扩展设施和融合的类型检查和计算,Coq是一个具有挑战性的目标,自动学习技术。我们提出了新的生成模型的学习和建议引理的Coq项目的名称。我们的模型基于多输入神经网络,是第一个利用来自Coq的词法分析器(lemma语句中的标记)、语法分析器(语法树)和内核(详细术语)的语法和语义信息进行命名的模型;关键的见解是,从详细术语中学习可以大大提高模型性能。我们在一个名为Roosterize的工具链中实现了我们的模型,并将其应用于一个来自Mathematical Components系列项目的大型代码库,该项目以其严格的编码约定而闻名。我们的研究结果表明,Roosterize在建议引理名称方面大大优于基线,突出了使用多输入模型和详细术语的重要性。
Coding conventions for naming, spacing, and other essentially stylistic properties are necessary for developers to effectively understand, review, and modify source code in large software projects. Consistent conventions in verification projects based on proof assistants, such as Coq, increase in importance as projects grow in size and scope. While conventions can be documented and enforced manually at high cost, emerging approaches automatically learn and suggest idiomatic names in Java-like languages by applying statistical language models on large code corpora. However, due to its powerful language extension facilities and fusion of type checking and computation, Coq is a challenging target for automated learning techniques. We present novel generation models for learning and suggesting lemma names for Coq projects. Our models, based on multi-input neural networks, are the first to leverage syntactic and semantic information from Coq ’s lexer (tokens in lemma statements), parser (syntax tree s), and kernel (elaborated terms) for naming; the key insight is that learning from elaborated terms can substantially boost model performance. We implemented our models in a toolchain, dubbed Roosterize, and applied it on a large corpus of code derived from the Mathematical Components family of projects, known for its stringent coding conventions. Our results show that Roosterize substantially outperforms baselines for suggesting lemma names, highlighting the importance of using multi-input models and elaborated terms.
DOI: 10.1023/a:1015761529444
发表时间: 2002-04-01
期刊: JOURNAL OF AUTOMATED REASONING
影响因子: --
作者:
Barendregt, H;Barendsen, E
通讯作者: Barendsen, E
DOI: 10.1007/s10817-018-9460-x
发表时间: 2018-06-01
期刊: JOURNAL OF AUTOMATED REASONING
影响因子: --
作者:
Doczkal, Christian;Smolka, Gert
通讯作者: Smolka, Gert
DOI: 10.1007/3-540-44404-1_7
发表时间: 2000-01-01
期刊: LOGIC FOR PROGRAMMING AND AUTOMATED REASONING, PROCEEDINGS
影响因子: --
作者:
Delahaye, D
通讯作者: Delahaye, D
DOI: 10.1145/3212695
发表时间: 2018-09-01
影响因子: 16.6
作者:
Allamanis, Miltiadis;Barr, Earl T.;Sutton, Charles
通讯作者: Sutton, Charles