White-box testing of big data analytics with complex user-defined functions

White-box testing of big data analytics with complex user-defined functions
复制标题

具有复杂的用户定义函数的大数据分析白盒测试

DOI:
10.1145/3338906.3338953
复制
发表时间:
2019
期刊:
ESEC/FSE 2019: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering
影响因子:
--
通讯作者:
Kim, Miryung
Kim, Miryung
中科院分区:
--
文献类型:
--
作者:
Gulzar, Muhammad Ali;Mardani, Shaghayegh;Musuvathi, Madanlal;Kim, Miryung

文献摘要

参考文献

被引文献

相似文献

数据密集型可伸缩计算(DISC)系统,如Google的MapReduce、Apache Hadoop和ApacheSpark,正被用于处理云中的海量数据。现代磁盘应用程序在详尽、自动的测试中提出了新的挑战,因为它们由数据流操作符组成,而且与SQL查询不同,复杂的用户定义函数(UDF)非常普遍。我们设计了一种新的白盒测试方法,称为BigTest,用于推理UDF的内部语义,并结合每个数据流和关系操作符创建的等价类。我们的评估表明,尽管存在超大规模的输入数据大小,但现实世界中的磁盘应用程序在测试覆盖范围方面往往存在严重的偏差和不足,使得34%的联合数据流和UDF(JDU)路径未进行测试。BigTest显示了将本地测试的数据大小最小化10^5到10^8个数量级的潜力,同时发现的手动注入故障比以前的方法多2倍。我们的实验表明,实际上只需要很少的数据记录(数量级为十)就可以实现与整个生产数据相同的JDU覆盖率。测试数据的减少还平均节省了194倍的CPU时间,证明了交互式、快速的本地测试对于大数据分析是可行的,从而消除了在海量生产数据上测试应用的需要。
Data-intensive scalable computing (DISC) systems such as Google’s MapReduce, Apache Hadoop, and Apache Spark are being leveraged to process massive quantities of data in the cloud. Modern DISC applications pose new challenges in exhaustive, automatic testing because they consist of dataflow operators, and complex user-defined functions (UDF) are prevalent unlike SQL queries. We design a new white-box testing approach, called BigTest to reason about the internal semantics of UDFs in tandem with the equivalence classes created by each dataflow and relational operator. Our evaluation shows that, despite ultra-large scale input data size, real world DISC applications are often significantly skewed and inadequate in terms of test coverage, leaving 34% of Joint Dataflow and UDF (JDU) paths untested. BigTest shows the potential to minimize data size for local testing by 10^5 to 10^8 orders of magnitude while revealing 2X more manually-injected faults than the previous approach. Our experiment shows that only few of the data records (order of tens) are actually required to achieve the same JDU coverage as the entire production data. The reduction in test data also provides CPU time saving of 194X on average, demonstrating that interactive and fast local testing is feasible for big data analytics, obviating the need to test applications on huge production data.
DOI: 10.1145/2025113.2025204
发表时间: 2011-09
期刊: --
影响因子: --
作者:
Christoph Csallner;L. Fegaras;Chengkai Li
通讯作者: Christoph Csallner;L. Fegaras;Chengkai Li
测试数据流程序运算符的属性
DOI: 10.1109/ase.2013.6693071
发表时间: 2013
期刊: 2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子: --
作者:
Zhihong Xu;Martin Hirzel;G. Rothermel;Kun
通讯作者: Kun
开放数据库连接
DOI: 10.1007/springerreference_64133
发表时间: 1999
期刊: Linux Journal
影响因子: --
作者:
Peter Harvey
通讯作者: Peter Harvey
符号测试和 DISSECT 符号评估系统
DOI: --
发表时间: 1977
影响因子: 7.4
作者:
W. Howden
通讯作者: W. Howden
SEDGE:数据流程序的符号示例数据生成
DOI: 10.1109/ase.2013.6693083
发表时间: 2013
期刊: 2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子: --
作者:
K. Li;Christoph Reichenbach;Y. Smaragdakis;Y. Diao;Christoph Csallner
通讯作者: Christoph Csallner