PerfDebug: Performance Debugging of Computation Skew in Dataflow Systems

PerfDebug: Performance Debugging of Computation Skew in Dataflow Systems
复制标题

DOI:
10.1145/3357223.3362727
复制
发表时间:
2019-11
期刊:
Proceedings of the ACM Symposium on Cloud Computing
影响因子:
--
通讯作者:
Jason Teoh;Muhammad Ali Gulzar;G. Xu;Miryung Kim
Jason Teoh;Muhammad Ali Gulzar;G. Xu;Miryung Kim
中科院分区:
其他
文献类型:
--
作者:
Jason Teoh;Muhammad Ali Gulzar;G. Xu;Miryung Kim

文献摘要

被引文献

相似文献

性能是大数据应用程序的关键因素,并且许多研究致力于优化这些应用程序。虽然先前的工作可以诊断和纠正数据偏斜,但计算问题的偏差 - 一小部分输入数据的计算成本异常高 - 在很大程度上被忽略了。计算通常发生在现实世界中的应用中,但没有工具可以供开发人员确定基本原因。为了使用户能够调试表现出偏斜的应用程序,我们开发了验尸绩效调试工具。 Perfdebug自动通过推理诸如工作执行时间,垃圾收集时间和序列化时间等性能指标的偏差来自动找到对大数据应用程序中此类异常的输入记录。 Perfdebug成功的关键是基于数据出处的技术,该技术可以计算和传播记录级计算潜伏期,以跟踪整个管道中异常昂贵的记录。最后,向用户提供了最大的延迟贡献的输入记录以进行错误修复。我们通过深入的案例研究评估了Perfdebug,并观察到诸如删除单个最昂贵的记录或简单代码重写之类的补救措施可以提高16倍的性能。
Performance is a key factor for big data applications, and much research has been devoted to optimizing these applications. While prior work can diagnose and correct data skew, the problem of computation skew---abnormally high computation costs for a small subset of input data---has been largely overlooked. Computation skew commonly occurs in real-world applications and yet no tool is available for developers to pinpoint underlying causes. To enable a user to debug applications that exhibit computation skew, we develop a post-mortem performance debugging tool. PerfDebug automatically finds input records responsible for such abnormalities in a big data application by reasoning about deviations in performance metrics such as job execution time, garbage collection time, and serialization time. The key to PerfDebug's success is a data provenance-based technique that computes and propagates record-level computation latency to keep track of abnormally expensive records throughout the pipeline. Finally, the input records that have the largest latency contributions are presented to the user for bug fixing. We evaluate PerfDebug via in-depth case studies and observe that remediation such as removing the single most expensive record or simple code rewrite can achieve up to 16X performance improvement.