Co-dependence Aware Fuzzing for Dataflow-Based Big Data Analytics

Co-dependence Aware Fuzzing for Dataflow-Based Big Data Analytics
复制标题

DOI:
10.1145/3611643.3616298
复制
发表时间:
2023-11
期刊:
Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering
影响因子:
--
通讯作者:
Ahmad Humayun;Miryung Kim;Muhammad Ali Gulzar
Ahmad Humayun;Miryung Kim;Muhammad Ali Gulzar
中科院分区:
其他
文献类型:
--
作者:
Ahmad Humayun;Miryung Kim;Muhammad Ali Gulzar

文献摘要

被引文献

相似文献

由于分析大数据的需求日益增长,数据密集型可扩展计算变得流行起来。例如,ApacheSpark和Hadoop允许开发人员使用用户定义的函数编写基于数据流的应用程序,以使用自定义逻辑处理数据。测试这类应用程序很困难。(1)这些应用程序通常接受多个数据集作为输入。(2)与SQL不同,这些数据集没有显式的模式,每个非结构化(或半结构化)数据集在运行时被分段和解析。(3)数据流运算符(例如,连接)在多个数据集的字段之间创建隐式相互依赖约束。一种高效有效的测试技术必须在行和列级别上分析多个数据集的不同区域之间的相互依赖关系,并在相互依赖的区域上联合编排输入突变。我们提出了DepFuzz来提高基于数据流的模糊测试大数据应用的有效性和效率。DepFuzz背后的关键洞察力是双重的。它跟踪哪些代码段在哪些数据集、哪些行和哪些列上操作。通过结合UDF的语义分析数据流运算符(例如,Join和groupByKey)的使用,DepFuzz生成测试数据,这些数据随后到达应用程序代码难以到达的区域。在现实世界的大数据应用程序中,DepFuzz发现的错误比Jazzer多3.4倍,语句覆盖率比Jazzer多29%,Jazzer是一款针对Java字节码的最先进的商业Fuzzer。它的性能优于以前的磁盘测试,因为它暴露了比更简单的输入格式化错误更深层次的语义错误,特别是当多个数据集通过数据流操作符进行复杂交互时。
Data-intensive scalable computing has become popular due to the increasing demands of analyzing big data. For example, Apache Spark and Hadoop allow developers to write dataflow-based applications with user-defined functions to process data with custom logic. Testing such applications is difficult. (1) These applications often take multiple datasets as input. (2) Unlike in SQL, there is no explicit schema for these datasets and each unstructured (or semi-structured) dataset is segmented and parsed at runtime. (3) Dataflow operators (e.g., join) create implicit co-dependence constraints between the fields of multiple datasets. An efficient and effective testing technique must analyze co-dependence among different regions of multiple datasets at the level of rows and columns and orchestrate input mutations jointly on co-dependent regions. We propose DepFuzz to increase the effectiveness and efficiency of fuzz testing dataflow-based big data applications. The key insight behind DepFuzz is twofold. It keeps track of which code segments operate on which datasets, which rows, and which columns. By analyzing the use of dataflow operators (e.g., join and groupByKey) in tandem with the semantics of UDFs, DepFuzz generates test data that subsequently reach hard-to-reach regions of the application code. In real-world big data applications, DepFuzz finds 3.4× more faults, achieving 29% more statement coverage in half the time as Jazzer’s, a state-of-the-art commercial fuzzer for Java bytecode. It outperforms prior DISC testing by exposing deeper semantic faults beyond simpler input formatting errors, especially when multiple datasets have complex interactions through dataflow operators.