On the Naturalness of Fuzzer-Generated Code

On the Naturalness of Fuzzer-Generated Code
复制标题

关于模糊器生成代码的自然性

DOI:
10.1145/3524842.3527972
复制
发表时间:
2022
期刊:
19th International Conference on Mining Software Repositories
影响因子:
--
通讯作者:
Hellendoorn, Vincent J.
Hellendoorn, Vincent J.
中科院分区:
--
文献类型:
--
作者:
Kambhamettu, Rajeswari Hita;Billos, John;Oluwaseun-Apo, Tomi;Gafford, Benjamin;Padhye, Rohan;Hellendoorn, Vincent J.

文献摘要

参考文献

相似文献

Csmith等编译器模糊工具通过从生成模型中随机抽样程序,在编译器中发现了许多错误。这些工具的成功通常归因于它们生成意外的角例输入的能力,而开发人员在手动测试过程中往往会忽略这些输入。同时,它们的混乱特性使得模糊生成的测试用例很难解释,这导致了输入简化工具的创建,如C-Reduced(针对C编译器错误)。在迄今为止无关的工作中,研究人员还表明,人类编写的软件往往具有相当的重复性,对语言模型来说是可以预测的。研究表明,开发人员故意编写更可预测的代码,而有错误的代码相对不可预测。在这项研究中,我们提出了一些自然的问题,即代码的这种高可预测性是否也适用于Fuzzer生成的代码,或许与直觉相反。也就是说,我们调查了模糊生成的编译器输入是否被构建在人类编写的代码上的语言模型认为是不可预测的,并出人意料地得出结论是不是的。相反,csmith fuzzer生成的程序在每个令牌的基础上比人类编写的C程序更可预测。此外,错误触发往往仍然比随机输入更具可预测性,而C-Reducts最小化工具并没有实质上增加这种可预测性。相反,我们发现,相对于Csmith自己的生成模型,错误触发输入是不可预测的。这是令人鼓舞的;我们的结果表明,在将可预测性指标纳入模糊化和约简工具本身方面,有希望的研究方向。
Compiler fuzzing tools such as Csmith have uncovered many bugs in compilers by randomly sampling programs from a generative model. The success of these tools is often attributed to their ability to generate unexpected corner case inputs that developers tend to overlook during manual testing. At the same time, their chaotic nature makes fuzzer-generated test cases notoriously hard to interpret, which has lead to the creation of input simplification tools such as C-Reduce (for C compiler bugs). In until now unrelated work, researchers have also shown that human-written software tends to be rather repetitive and predictable to language models. Studies show that developers deliberately write more predictable code, whereas code with bugs is relatively unpredictable. In this study, we ask the natural questions of whether this high predictability property of code also, and perhaps counter-intuitively, applies to fuzzer-generated code. That is, we investigate whether fuzzer-generated compiler inputs are deemed unpredictable by a language model built on human-written code and surprisingly conclude that it is not. To the contrary, Csmith fuzzer-generated programs aremorepredictable on a per-token basis than human-written C programs. Furthermore, bug-triggering tended to be more predictable still than random inputs, and the C-Reduce minimization tool did not substantially increase this predictability. Rather, we find that bug-triggering inputs are unpredictable relative toCsmith's owngenerative model. This is encouraging; our results suggest promising research directions on incorporating predictability metrics in the fuzzing and reduction tools themselves.
DeepTC-Enhancer:提高自动生成测试的可读性
DOI: 10.1145/3324884.3416622
发表时间: 2020
期刊: 2020 35th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子: --
作者:
Devjeet Roy;Ziyi Zhang;Maggie Ma;Venera Arnaoudova;Annibale Panichella;Sebastiano Panichella;Danielle Gonzalez;Mehdi Mirakhorli
通讯作者: Mehdi Mirakhorli