DeepLocalize: Fault Localization for Deep Neural Networks

DeepLocalize: Fault Localization for Deep Neural Networks
复制标题

DOI:
10.1109/icse43902.2021.00034
复制
发表时间:
2021-03
期刊:
2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Mohammad Wardat;Wei Le;Hridesh Rajan
Mohammad Wardat;Wei Le;Hridesh Rajan
中科院分区:
其他
文献类型:
--
作者:
Mohammad Wardat;Wei Le;Hridesh Rajan

文献摘要

相似文献

深度神经网络(DNN)正在成为大多数软件系统不可或缺的一部分。以前的工作表明DNN有bug。不幸的是,现有的调试技术不支持本地化DNN错误,因为缺乏对模型行为的理解。整个DNN模型显示为一个黑盒子。为了解决这些问题,我们提出了一种方法和工具,可以自动确定模型是否存在错误,并确定DNN错误的根本原因。我们的关键见解是,可以分析层之间传播的值的历史趋势,以识别故障,并定位故障。为此,我们首先启用深度学习应用程序的动态分析:将其转换为命令式表示,或者使用回调机制。这两种机制都允许我们插入探针,以便在DNN在训练数据上进行训练时对DNN产生的轨迹进行动态分析。然后,我们对迹线进行动态分析,以识别导致错误的故障层或超参数。我们提出了一种算法,通过捕获任何数值错误来识别根本原因,并在训练过程中监控模型,并找到DNN结果中每个层/参数的相关性。我们从Stack Overflow和GitHub收集了一个基准测试,其中包含40个错误模型和补丁,这些模型和补丁包含深度学习应用程序中的真实的错误。我们的基准可以用来评估自动调试工具和修复技术。我们已经使用这个DNN bug和补丁基准测试对我们的方法进行了评估,结果表明我们的方法比Keras库中使用的现有调试方法更有效。对于34/40的情况,我们的方法能够检测到错误,而Keras提供的最佳调试方法检测到32/40的错误。我们的方法能够定位21/40个错误,而Keras没有定位任何错误。
Deep Neural Networks (DNNs) are becoming an integral part of most software systems. Previous work has shown that DNNs have bugs. Unfortunately, existing debugging techniques don't support localizing DNN bugs because of the lack of understanding of model behaviors. The entire DNN model appears as a black box. To address these problems, we propose an approach and a tool that automatically determines whether the model is buggy or not, and identifies the root causes for DNN errors. Our key insight is that historic trends in values propagated between layers can be analyzed to identify faults, and also localize faults. To that end, we first enable dynamic analysis of deep learning applications: by converting it into an imperative representation and alternatively using a callback mechanism. Both mechanisms allows us to insert probes that enable dynamic analysis over the traces produced by the DNN while it is being trained on the training data. We then conduct dynamic analysis over the traces to identify the faulty layer or hyperparameter that causes the error. We propose an algorithm for identifying root causes by capturing any numerical error and monitoring the model during training and finding the relevance of every layer/parameter on the DNN outcome. We have collected a benchmark containing 40 buggy models and patches that contain real errors in deep learning applications from Stack Overflow and GitHub. Our benchmark can be used to evaluate automated debugging tools and repair techniques. We have evaluated our approach using this DNN bug-and-patch benchmark, and the results showed that our approach is much more effective than the existing debugging approach used in the state-of-the-practice Keras library. For 34/40 cases, our approach was able to detect faults whereas the best debugging approach provided by Keras detected 32/40 faults. Our approach was able to localize 21/40 bugs whereas Keras did not localize any faults.