BigTest: A Symbolic Execution Based Systematic Test Generation Tool for Apache Spark

BigTest: A Symbolic Execution Based Systematic Test Generation Tool for Apache Spark
复制标题

BigTest:基于符号执行的 Apache Spark 系统测试生成工具

DOI:
10.1145/3377812.3382145
复制
发表时间:
2020
期刊:
Proceedings of 42nd IEEE/ACM International Conference on Software Engineering
影响因子:
--
通讯作者:
Kim, Miryung
Kim, Miryung
中科院分区:
--
文献类型:
--
作者:
Gulzar, Muhammad Ali;Musuvathi, Madanlal;Kim, Miryung

文献摘要

参考文献

被引文献

相似文献

数据密集型可扩展计算(DISC)系统,如Google的MapReduce,Apache Hadoop和Apache Spark,在许多生产服务中很流行。尽管DISC应用程序很受欢迎,但由于缺乏详尽和自动化的测试,其质量受到影响。目前测试DISC应用程序的实践仅限于使用整个输入数据集的一个小的随机样本,这仅仅暴露了任何程序错误。与SQL查询不同的是,测试DISC应用程序有新的挑战,由于一个组合的关系操作符,和用户定义的功能(UDF),可以是任意长和复杂的。为了解决这个问题,我们展示了一个新的白盒测试框架,称为BigTest,采用Apache Spark程序作为输入,并自动生成合成,具体的数据进行有效和高效的测试。BigTest将UDF的符号执行与XML和关系运算符的逻辑规范相结合,以探索DISC应用程序中的所有路径。我们的实验表明,BigTest能够生成测试数据,这些数据可以比整个数据集显示多达2倍的故障,而测试时间减少了194倍。我们在基于Java的命令行工具中使用预编译的二进制jar实现了BigTest。它公开了一个配置文件,用户可以在其中编辑首选项,包括目标程序的路径、循环探索的上限以及定理求解器的选择。BigTest的演示视频可在https://youtu.be/OeHhoKiDYso上获得,BigTest可在https://github.com/maligulzar/BigTest上获得。
Data-intensive scalable computing (DISC) systems such as Google's MapReduce, Apache Hadoop, and Apache Spark are prevalent in many production services. Despite their popularity, the quality of DISC applications suffers due to a lack of exhaustive and automated testing. Current practices of testing DISC applications are limited to using a small random sample of the entire input dataset which merely exposes any program faults. Unlike SQL queries, testing DISC applications has new challenges due to a composition of both dataflow and relational operators, and user-defined functions (UDF) that could be arbitrarily long and complex.To address this problem, we demonstrate a new white-box testing framework called BigTest that takes an Apache Spark program as input and automatically generates synthetic, concrete data for effective and efficient testing. BigTest combines the symbolic execution of UDFs with the logical specifications of dataflow and relational operators to explore all paths in a DISC application. Our experiments show that BigTest is capable of generating test data that can reveal up to 2X more faults than the entire data set with 194X less testing time. We implement BigTest in a Java-based command line tool with a pre-compile binary jar. It exposes a configuration file in which a user can edit preferences, including the path of a target program, the upper bound of loop exploration, and a choice of theorem solver. The demonstration video of BigTest is available at https://youtu.be/OeHhoKiDYso and BigTest is available at https://github.com/maligulzar/BigTest.
具有复杂的用户定义函数的大数据分析白盒测试
DOI: 10.1145/3338906.3338953
发表时间: 2019
期刊: ESEC/FSE 2019: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering
影响因子: --
作者:
Gulzar, Muhammad Ali;Mardani, Shaghayegh;Musuvathi, Madanlal;Kim, Miryung
通讯作者: Kim, Miryung
SEDGE:数据流程序的符号示例数据生成
DOI: 10.1109/ase.2013.6693083
发表时间: 2013
期刊: 2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子: --
作者:
K. Li;Christoph Reichenbach;Y. Smaragdakis;Y. Diao;Christoph Csallner
通讯作者: Christoph Csallner