SCMiner: Localizing System-Level Concurrency Faults from Large System Call Traces

SCMiner: Localizing System-Level Concurrency Faults from Large System Call Traces
复制标题

DOI:
10.1109/ase.2019.00055
复制
发表时间:
2019-11
期刊:
2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
T. S. Zaman;Xue Han;Tingting Yu
T. S. Zaman;Xue Han;Tingting Yu
中科院分区:
其他
文献类型:
--
作者:
T. S. Zaman;Xue Han;Tingting Yu

文献摘要

相似文献

生产中发生的定位并发故障很难,因为(1)详细的字段数据(例如用户输入,文件内容和交织时间表)可能无法用于开发人员重现故障; (2)假设使用多个失败的执行来使用现有技术定位故障通常是不切实际的; (3)在给定有限的运行时数据中搜索应用程序中的错误位置是一项挑战; (4)系统级别的并发故障通常涉及多个进程或事件处理程序(例如,软件信号),无法通过现有工具来诊断内部过程(线程级别)故障的现有工具来处理。为了解决这些问题,我们提出了SCMiner,这是一种实用的在线错误诊断工具,可帮助开发人员根据默认系统审核工具收集的日志来了解系统级并发故障如何发生。 SCMiner实现在线错误诊断,以消除对离线错误复制的需求。 SCMiner不需要在生产系统上的代码仪器或依靠多个失败执行的可用性的假设。具体而言,在收集系统调用轨迹后,SCMiner使用数据挖掘和统计异常检测技术来识别诱导失败的系统调用序列。然后,它将每个异常序列映射到特定的应用函数。我们已经对19个现实世界的基准进行了一项实证研究。结果表明,SCMiner在定位系统级并发故障方面既有效又有效。
Localizing concurrency faults that occur in production is hard because, (1) detailed field data, such as user input, file content and interleaving schedule, may not be available to developers to reproduce the failure; (2) it is often impractical to assume the availability of multiple failing executions to localize the faults using existing techniques; (3) it is challenging to search for buggy locations in an application given limited runtime data; and, (4) concurrency failures at the system level often involve multiple processes or event handlers (e.g., software signals), which can not be handled by existing tools for diagnosing intra-process(thread-level) failures. To address these problems, we present SCMiner, a practical online bug diagnosis tool to help developers understand how a system-level concurrency fault happens based on the logs collected by the default system audit tools. SCMiner achieves online bug diagnosis to obviate the need for offline bug reproduction. SCMiner does not require code instrumentation on the production system or rely on the assumption of the availability of multiple failing executions. Specifically, after the system call traces are collected, SCMiner uses data mining and statistical anomaly detection techniques to identify the failure-inducing system call sequences. It then maps each abnormal sequence to specific application functions. We have conducted an empirical study on 19 real-world benchmarks. The results show that SCMiner is both effective and efficient at localizing system-level concurrency faults.