Test and Yield Loss Reduction of AI and Deep Learning Accelerators

Test and Yield Loss Reduction of AI and Deep Learning Accelerators
复制标题

人工智能和深度学习加速器的测试和良率损失降低

DOI:
--
复制
发表时间:
2021
影响因子:
2.9
通讯作者:
Ujjwal Guin
Ujjwal Guin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Mehdi Sadi;Ujjwal Guin

文献摘要

被引文献

相似文献

随着数据驱动分析成为主流,全球对专用人工智能 (AI) 和深度学习加速器芯片的需求正在飙升。这些加速器采用密集封装的处理元件 (PE) 设计,特别容易受到先进半导体工艺节点中常见的制造缺陷和功能故障的影响,从而导致显着的产量损失。在这项工作中,我们演示了一种应用驱动的方法,对 AI 加速器芯片进行分级,并通过将加速器 PE 中的电路故障与目标 AI 工作负载的所需精度相关联来减少良率损失。我们利用经过训练的深度学习模型固有的容错特性和选择性停用故障 PE 的策略来开发所提出的良率损失减少和测试方法。推导出故障位置、故障率和人工智能任务准确性之间的分析关系,以决定加速器芯片是否可以通过最终的良率测试。针对 PE 的乘法和累加单元,提出了良率损失减少感知的故障隔离、ATPG 和测试流程。通过广泛使用的人工智能/深度学习基准获得的结果表明,加速器可以在 PE 阵列中维持 5% 的故障率,同时精度损失不到 1%,从而实现这些芯片的产品分档和良率损失。
With data-driven analytics becoming mainstream, the global demand for dedicated artificial intelligence (AI) and deep learning accelerator chips is soaring. These accelerators, designed with densely packed processing elements (PE), are especially vulnerable to the manufacturing defects and functional faults common in the advanced semiconductor process nodes resulting in significant yield loss. In this work, we demonstrate an application-driven methodology of binning the AI accelerator chips, and yield loss reduction by correlating the circuit faults in the PEs of the accelerator with the desired accuracy of the target AI workload. We exploit the inherent fault tolerance features of trained deep learning models and a strategy of selective deactivation of faulty PEs to develop the presented yield loss reduction and test methodology. An analytical relationship is derived between fault location, fault rate, and the AI task’s accuracy for deciding if the accelerator chip can pass the final yield test. A yield-loss reduction-aware fault isolation, ATPG, and test flow are presented for the multiply and accumulate units of the PEs. Results obtained with widely used AI/deep learning benchmarks demonstrate that the accelerators can sustain 5% fault rate in PE arrays while suffering from less than 1% accuracy loss, thus enabling product binning and yield loss reduction of these chips.