SEDGE: Symbolic example data generation for dataflow programs

SEDGE: Symbolic example data generation for dataflow programs
复制标题

SEDGE:数据流程序的符号示例数据生成

DOI:
10.1109/ase.2013.6693083
复制
发表时间:
2013
期刊:
2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Christoph Csallner
Christoph Csallner
中科院分区:
--
文献类型:
--
作者:
K. Li;Christoph Reichenbach;Y. Smaragdakis;Y. Diao;Christoph Csallner

文献摘要

被引文献

相似文献

详尽的数据流程(尤其是MAPREDUCE)的程序已成为一个重要的挑战,这表明了在猪平台上锻炼运算符的小示例数据集,这些数据集用于生成Hadoop Map-Reduce程序技术试图涵盖所有操作员使用的案例,实际上,我们的SEDGE系统会解决这些完整性问题:对于每个数据流操作员,我们都会产生数据瞄准涵盖数据流程中出现的所有情况(例如,通过和失败的过滤器)依靠将程序转换为符号约束,并使用符号推理引擎(强大的SMT求解器)来解决约束数据作为解决方案过程,类似于动态符号(又称“ Concolic”)的传统编程语言,适用于数据流域的独特功能第三方基准测试,莎士的覆盖范围比过去的技术更高,在20个鸽子基准中,有7个,在11个SDSS基准中,有7个,并且(在其余基准中,我们还表明我们对高级目标的目标)。 DataFlow语言回报:对于复杂的程序,在生成的MAP-REDUCE代码级别(而不是原始DataFlow程序)的最新动态符号执行需要更多测试案件或取得的覆盖范围要比我们的方法低得多。
Exhaustive, automatic testing of dataflow (esp. mapreduce) programs has emerged as an important challenge. Past work demonstrated effective ways to generate small example data sets that exercise operators in the Pig platform, used to generate Hadoop map-reduce programs. Although such prior techniques attempt to cover all cases of operator use, in practice they often fail. Our SEDGE system addresses these completeness problems: for every dataflow operator, we produce data aiming to cover all cases that arise in the dataflow program (e.g., both passing and failing a filter). SEDGE relies on transforming the program into symbolic constraints, and solving the constraints using a symbolic reasoning engine (a powerful SMT solver), while using input data as concrete aids in the solution process. The approach resembles dynamic-symbolic (a.k.a. “concolic”) execution in a conventional programming language, adapted to the unique features of the dataflow domain. In third-party benchmarks, SEDGE achieves higher coverage than past techniques for 5 out of 20 PigMix benchmarks and 7 out of 11 SDSS benchmarks and (with equal coverage for the rest of the benchmarks). We also show that our targeting of the high-level dataflow language pays off: for complex programs, state-of-the-art dynamic-symbolic execution at the level of the generated map-reduce code (instead of the original dataflow program) requires many more test cases or achieves much lower coverage than our approach.