Monitoring the Health of Emerging Neural Network Accelerators with Cost-effective Concurrent Test

Monitoring the Health of Emerging Neural Network Accelerators with Cost-effective Concurrent Test
复制标题

DOI:
10.1109/dac18072.2020.9218675
复制
发表时间:
2020-07
期刊:
2020 57th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Qi Liu;Tao Liu;Zihao Liu;Wujie Wen;Chengmo Yang
Qi Liu;Tao Liu;Zihao Liu;Wujie Wen;Chengmo Yang
中科院分区:
其他
文献类型:
--
作者:
Qi Liu;Tao Liu;Zihao Liu;Wujie Wen;Chengmo Yang

文献摘要

相似文献

基于ReRAM的神经网络加速器是处理内存和计算密集型深度学习工作负载的一种有前途的解决方案。然而,它遭受独特的设备错误。这些错误可能在运行期间累积到大量水平,并导致显着的准确性下降。在应用任何适当的修复机制之前,实时获取其故障状态至关重要。然而,校准这样的统计信息是不平凡的,因为需要大量的测试模式,长测试时间,和高测试覆盖率,考虑到复杂的错误可能出现在百万到十亿的权重参数。在本文中,我们利用角落数据的概念,可以显着混淆神经网络模型的决策,以及训练算法,只生成一小部分测试模式,调整为对不同程度的误差积累和准确性损失敏感。实验结果表明,该方法能够快速、准确地报告正在运行的加速器的故障状态,在检测效率和成本方面均优于现有的解决方案
ReRAM-based neural network accelerator is a promising solution to handle the memory-and computation-intensive deep learning workloads. However, it suffers from unique device errors. These errors can accumulate to massive levels during the run time and cause significant accuracy drop. It is crucial to obtain its fault status in real-time before any proper repair mechanism can be applied. However, calibrating such statistical information is non-trivial because of the need of a large number of test patterns, long test time, and high test coverage considering that complex errors may appear in million-to-billion weight parameters. In this paper, we leverage the concept of comer data that can significantly confuse the decision making of neural network model, as well as the training algorithm, to generate only a small set of test patterns that is tuned to be sensitive to different levels of error accumulation and accuracy loss. Experimental results show that our method can quickly and correctly report the fault status of a running accelerator, outperforming existing solutions in both detection efficiency and cost