DeepPerform: An Efficient Approach for Performance Testing of Resource-Constrained Neural Networks

DeepPerform: An Efficient Approach for Performance Testing of Resource-Constrained Neural Networks
复制标题

DOI:
10.1145/3551349.3561158
复制
发表时间:
2022-10
期刊:
Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering
影响因子:
--
通讯作者:
Simin Chen;Mirazul Haque;Cong Liu;Wei Yang
Simin Chen;Mirazul Haque;Cong Liu;Wei Yang
中科院分区:
其他
文献类型:
--
作者:
Simin Chen;Mirazul Haque;Cong Liu;Wei Yang

文献摘要

相似文献

如今,越来越多的自适应深度神经网络(AdNns)被用于资源受限的嵌入式设备。我们观察到,与传统软件类似,AdNN中存在冗余计算,导致性能显著下降。性能下降依赖于输入,称为依赖于输入的性能瓶颈(IDPB)。为了确保AdNN满足资源受限应用的性能要求,必须进行性能测试来检测AdNN中的IDPB。现有的神经网络测试方法主要关注正确性测试,而不涉及性能测试。为了填补这一空白,我们提出了DeepPerform,一种可扩展的方法来生成测试样本来检测AdNN中的IDPB。我们首先演示了如何将生成检测IDPB的性能测试样本的问题描述为一个优化问题。然后,我们展示了DeepPerform如何通过学习和估计AdNns的计算消耗的分布来有效地处理优化问题。我们在三个广泛使用的数据集和五个流行的AdNN模型上对DeepPerform进行了评估。结果显示,DeepPerform生成的测试样本会导致更严重的性能下降(Flops:增加高达552%)。此外,在生成测试输入方面,DeepPerform比基线方法更有效(运行时开销仅为6-10毫秒)。
Today, an increasing number of Adaptive Deep Neural Networks (AdNNs) are being used on resource-constrained embedded devices. We observe that, similar to traditional software, redundant computation exists in AdNNs, resulting in considerable performance degradation. The performance degradation is dependent on the input and is referred to as input-dependent performance bottlenecks (IDPBs). To ensure an AdNN satisfies the performance requirements of resource-constrained applications, it is essential to conduct performance testing to detect IDPBs in the AdNN. Existing neural network testing methods are primarily concerned with correctness testing, which does not involve performance testing. To fill this gap, we propose DeepPerform, a scalable approach to generate test samples to detect the IDPBs in AdNNs. We first demonstrate how the problem of generating performance test samples detecting IDPBs can be formulated as an optimization problem. Following that, we demonstrate how DeepPerform efficiently handles the optimization problem by learning and estimating the distribution of AdNNs’ computational consumption. We evaluate DeepPerform on three widely used datasets against five popular AdNN models. The results show that DeepPerform generates test samples that cause more severe performance degradation (FLOPs: increase up to 552%). Furthermore, DeepPerform is substantially more efficient than the baseline methods in generating test inputs (runtime overhead: only 6–10 milliseconds).