UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating Fuzzers

UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating Fuzzers
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Yuwei Li;S. Ji;Yuan Chen;Sizhuang Liang;Wei-Han Lee;Yueyao Chen;Chenyang Lyu-;Chunming Wu;R. Be
Yuwei Li;S. Ji;Yuan Chen;Sizhuang Liang;Wei-Han Lee;Yueyao Chen;Chenyang Lyu-;Chunming Wu;R. Be
中科院分区:
其他
文献类型:
--
作者:
Yuwei Li;S. Ji;Yuan Chen;Sizhuang Liang;Wei-Han Lee;Yueyao Chen;Chenyang Lyu-;Chunming Wu;R. Be

文献摘要

被引文献

相似文献

文献中已经提出了一系列模糊测试工具(模糊器),旨在有效且高效地检测软件漏洞。然而,迄今为止,由于基准、性能指标和/或评估环境的不一致,比较模糊器仍然具有挑战性,这掩盖了有用的见解,从而阻碍了有前途的模糊原语的发现。在本文中,我们设计和开发了 UNIFUZZ,这是一个开源和指标驱动的平台,用于以全面和定量的方式评估模糊器。具体来说,UNIFUZZ 迄今为止已包含 35 个可用的模糊器、20 个实际程序的基准以及六类性能指标。我们首先系统地研究了现有模糊器的可用性,发现并修复了一些缺陷,并将其集成到 UNIFUZZ 中。基于这项研究,我们提出了一系列实用的性能指标,从六个互补的角度来评估模糊器。使用 UNIFUZZ,我们对几个著名的模糊器进行了深入评估,包括 AFL [1]、AFLFast [2]、Angora [3]、honggfuzz [4]、MOPT [5]、QSYM [6]、T-Fuzz [7] 和 VUzzer64 [8]。我们发现,在所有目标程序中,它们都没有优于其他程序,并且使用单一指标来评估模糊器的性能可能会导致单方面的结论,这证明了综合指标的重要性。此外,我们还识别并研究了以前被忽视的可能显着影响模糊器性能的因素,包括仪器方法和崩溃分析工具。我们的实证结果表明,它们对于模糊器的评估至关重要。我们希望我们的研究结果能够揭示可靠的模糊评估,以便我们能够发现有前途的模糊原语,以有效促进未来的模糊器设计。
A flurry of fuzzing tools (fuzzers) have been proposed in the literature, aiming at detecting software vulnerabilities effectively and efficiently. To date, it is however still challenging to compare fuzzers due to the inconsistency of the benchmarks, performance metrics, and/or environments for evaluation, which buries the useful insights and thus impedes the discovery of promising fuzzing primitives. In this paper, we design and develop UNIFUZZ, an open-source and metrics-driven platform for assessing fuzzers in a comprehensive and quantitative manner. Specifically, UNIFUZZ to date has incorporated 35 usable fuzzers, a benchmark of 20 real-world programs, and six categories of performance metrics. We first systematically study the usability of existing fuzzers, find and fix a number of flaws, and integrate them into UNIFUZZ. Based on the study, we propose a collection of pragmatic performance metrics to evaluate fuzzers from six complementary perspectives. Using UNIFUZZ, we conduct in-depth evaluations of several prominent fuzzers including AFL [1], AFLFast [2], Angora [3], Honggfuzz [4], MOPT [5], QSYM [6], T-Fuzz [7] and VUzzer64 [8]. We find that none of them outperforms the others across all the target programs, and that using a single metric to assess the performance of a fuzzer may lead to unilateral conclusions, which demonstrates the significance of comprehensive metrics. Moreover, we identify and investigate previously overlooked factors that may significantly affect a fuzzer's performance, including instrumentation methods and crash analysis tools. Our empirical results show that they are critical to the evaluation of a fuzzer. We hope that our findings can shed light on reliable fuzzing evaluation, so that we can discover promising fuzzing primitives to effectively facilitate fuzzer designs in the future.