The Relevance of Classic Fuzz Testing: Have We Solved This One?

The Relevance of Classic Fuzz Testing: Have We Solved This One?
复制标题

DOI:
10.1109/tse.2020.3047766
复制
发表时间:
2020-08
影响因子:
7.4
通讯作者:
B. Miller;Mengxiao Zhang;E. Heymann
B. Miller;Mengxiao Zhang;E. Heymann
中科院分区:
计算机科学1区
文献类型:
--
作者:
B. Miller;Mengxiao Zhang;E. Heymann

文献摘要

相似文献

随着模糊测试诞生30周年,面对模糊测试技术和工具令人难以置信的进步,问题来了,经典的、基本的模糊技术是否仍然有用和适用?在这一传统中,我们更新了基本的模糊工具和测试脚本,并将它们应用于Linux、FreeBSD和MacOS上的大量Unix实用程序。和以前一样,我们的失败标准是程序是否崩溃或挂起。我们发现Linux上74个实用程序中有9个崩溃或挂起,FreeBSD上78个实用程序中有15个,MacOS上76个实用程序中有12个。三个平台上共有24个不同的公用事业失败。我们注意到,这些故障率比我们在1995年、2000年和2006年对命令行实用程序可靠性的研究要高一些。在基本的模糊传统中,我们调试每个失败的实用程序,并对失败的原因进行分类。经典类型的失败,如指针和数组错误以及未检查返回代码,在当前的结果中仍然广泛存在。此外,我们还发现出现了一些新的故障类别。我们提供了这些失败的例子来说明导致它们发生的编程实践。顺便说一句,我们测试了现代编程语言(Rust)中可用的有限数量的实用程序,发现它们的可靠性并不比标准工具好。
As fuzz testing has passed its 30th anniversary, and in the face of the incredible progress in fuzz testing techniques and tools, the question arises if the classic, basic fuzz technique is still useful and applicable? In that tradition, we have updated the basic fuzz tools and testing scripts and applied them to a large collection of Unix utilities on Linux, FreeBSD, and MacOS. As before, our failure criteria was whether the program crashed or hung. We found that 9 crash or hang out of 74 utilities on Linux, 15 out of 78 utilities on FreeBSD, and 12 out of 76 utilities on MacOS. A total of 24 different utilities failed across the three platforms. We note that these failure rates are somewhat higher than our in previous 1995, 2000, and 2006 studies of the reliability of command line utilities. In the basic fuzz tradition, we debugged each failed utility and categorized the causes the failures. Classic categories of failures, such as pointer and array errors and not checking return codes, were still broadly present in the current results. In addition, we found a couple of new categories of failures appearing. We present examples of these failures to illustrate the programming practices that allowed them to happen. As a side note, we tested the limited number of utilities available in a modern programming language (Rust) and found them to be of no better reliability than the standard ones.