EAGLE: Creating Equivalent Graphs to Test Deep Learning Libraries

EAGLE: Creating Equivalent Graphs to Test Deep Learning Libraries
复制标题

DOI:
10.1145/3510003.3510165
复制
发表时间:
2022-05
期刊:
2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Jiannan Wang;Thibaud Lutellier;Shangshu Qian;H. Pham;Lin Tan
Jiannan Wang;Thibaud Lutellier;Shangshu Qian;H. Pham;Lin Tan
中科院分区:
其他
文献类型:
--
作者:
Jiannan Wang;Thibaud Lutellier;Shangshu Qian;H. Pham;Lin Tan

文献摘要

相似文献

测试深度学习(DL)软件至关重要且具有挑战性。最近的方法使用差分测试来交叉检查不同库中相同功能的实现对。这种方法需要两个实现相同功能的DL库,这通常不可用。此外,他们依靠一个高级库KERA,该图书馆在所有受支持的DL库中都实现了缺少功能,该库非常昂贵,因此不再维护。为了解决这个问题,我们建议通过使用等效图测试单个DL实现(例如,单个DL库),将Eagle(一种新技术)使用不同维度的新技术。等效图使用不同的应用程序编程接口(API),数据类型或优化来实现相同的功能。理由是,在单个DL实现上执行的两个等效图应在给定相同输入的情况下产生相同的输出。具体而言,我们设计了16个新的DL等价规则,并提出了一种技术,Eagle,(1)使用这些等价规则来构建等价图的具体对,(2)交叉检查这些等效图的输出以检测一个在一个中的不一致错误DL库。我们对两个广泛使用的DL库的评估,即张量和Pytorch,这表明Eagle检测到25个错误(张量为18个,pytorch中有7个),包括13个以前未知的错误。
Testing deep learning (DL) software is crucial and challenging. Recent approaches use differential testing to cross-check pairs of implementations of the same functionality across different libraries. Such approaches require two DL libraries implementing the same functionality, which is often unavailable. In addition, they rely on a high-level library, Keras, that implements missing functionality in all supported DL libraries, which is prohibitively expensive and thus no longer maintained. To address this issue, we propose EAGLE, a new technique that uses differential testing in a different dimension, by using equivalent graphs to test a single DL implementation (e.g., a single DL library). Equivalent graphs use different Application Programming Interfaces (APIs), data types, or optimizations to achieve the same functionality. The rationale is that two equivalent graphs executed on a single DL implementation should produce identical output given the same input. Specifically, we design 16 new DL equivalence rules and propose a technique, EAGLE, that (1) uses these equivalence rules to build concrete pairs of equivalent graphs and (2) cross-checks the output of these equivalent graphs to detect inconsistency bugs in a DL library. Our evaluation on two widely-used DL libraries, i.e., Tensor Flow and PyTorch, shows that EAGLE detects 25 bugs (18 in Tensor Flow and 7 in PyTorch), including 13 previously unknown bugs.