Learning to blame: localizing novice type errors with data-driven diagnosis

Learning to blame: localizing novice type errors with data-driven diagnosis
复制标题

学会责备:通过数据驱动的诊断来定位新手类型错误

DOI:
--
复制
发表时间:
2017
期刊:
Proc. ACM Program. Lang.
影响因子:
--
通讯作者:
Ranjit Jhala
Ranjit Jhala
中科院分区:
--
文献类型:
--
作者:
Eric L. Seidel;Huma Sibghat;Kamalika Chaudhuri;Westley Weimer;Ranjit Jhala

文献摘要

被引文献

相似文献

在具有全局类型推断的语言中,本地化类型错误是一项挑战,因为类型检查器必须对程序员的意图做出假设。介绍了一种基于监督学习的数据驱动错误定位方法NATE。Nate分析了大量的训练数据--输入错误的程序对及其“固定”版本--以自动学习最有可能在哪里发现错误的模型。给出一个新的错误类型的程序,Nate执行该模型以生成一个按可能性排序的潜在责任分配列表。我们通过将Nate的精确度与从入门编程课程的两个实例中提取的5000多个错误类型的OCaml程序集上的最新技术进行比较来评估Nate。我们发现,当考虑到排名靠前的指责分配时,Nate的数据驱动模型能够正确预测应该被更改的精确子表达式的概率为72%,比OCaml高出28个百分点,比最先进的SHErrLoc工具高出16个百分点。此外,当我们考虑前两个位置时,Nate的准确率超过85%,如果我们考虑前三个位置,Nate的准确率达到91%。
Localizing type errors is challenging in languages with global type inference, as the type checker must make assumptions about what the programmer intended to do. We introduce Nate, a data-driven approach to error localization based on supervised learning. Nate analyzes a large corpus of training data -- pairs of ill-typed programs and their "fixed" versions -- to automatically learn a model of where the error is most likely to be found. Given a new ill-typed program, Nate executes the model to generate a list of potential blame assignments ranked by likelihood. We evaluate Nate by comparing its precision to the state of the art on a set of over 5,000 ill-typed OCaml programs drawn from two instances of an introductory programming course. We show that when the top-ranked blame assignment is considered, Nate's data-driven model is able to correctly predict the exact sub-expression that should be changed 72% of the time, 28 points higher than OCaml and 16 points higher than the state-of-the-art SHErrLoc tool. Furthermore, Nate's accuracy surpasses 85% when we consider the top two locations and reaches 91% if we consider the top three.