Designing Efficient and High-Performance AI Accelerators With Customized STT-MRAM

Designing Efficient and High-Performance AI Accelerators With Customized STT-MRAM
复制标题

DOI:
10.1109/tvlsi.2021.3105958
复制
发表时间:
2021-10-01
影响因子:
2.8
通讯作者:
Sadi, Mehdi
Sadi, Mehdi
中科院分区:
工程技术2区
文献类型:
--
作者:
Mishty, Kaniz;Sadi, Mehdi

文献摘要

被引文献

相似文献

我们展示了高效和高性能的人工智能(AI)/深度学习加速器的设计,该加速器具有定制的自旋转移矩(STT)-MRAM(STT-MRAM)和可重新配置的内核。基于模型驱动的详细设计空间探索,我们提出了一个创新的暂存器辅助片上STT-MRAM为基础的高性能加速器的缓冲系统的设计方法。使用AI模型权重和激活图的存储器占用时间的解析导出的表达式,STT-MRAM的易失性利用热稳定性因子的工艺和温度变化感知缩放来调整,以优化STT-MRAM的保留时间、能量、读/写延迟和面积。通过对14 nm技术中AI工作负载和加速器实现的分析,我们验证了我们的STT-MRAM(STT-AI)AI加速器的有效性。与基于SRAM的实现相比,STT-AI加速器在等精度下实现了75%的面积和3%的功耗节省。此外,在宽松的误码率和可忽略不计的AI精度权衡下,所设计的STT-AI Ultra加速器比常规的基于SRAM的加速器分别节省了75.4%和3.5%的面积和功耗。
We demonstrate the design of efficient and high-performance artificial intelligence (AI)/deep learning accelerators with customized spin transfer torque (STT)-MRAM (STT-MRAM) and a reconfigurable core. Based on model-driven detailed design space exploration, we present the design methodology of an innovative scratchpad-assisted on-chip STT-MRAM-based buffer system for high-performance accelerators. Using analytically derived expression of memory occupancy time of AI model weights and activation maps, the volatility of STT-MRAM is adjusted with process and temperature variation aware scaling of thermal stability factor to optimize the retention time, energy, read/write latency, and area of STT-MRAM. From the analysis of AI workloads and accelerator implementation in 14-nm technology, we verify the efficacy of our AI accelerator with STT-MRAM (STT-AI). Compared to an SRAM-based implementation, the STT-AI accelerator achieves 75% area and 3% power savings at isoaccuracy. Furthermore, with a relaxed bit error rate and negligible AI accuracy tradeoff, the designed STT-AI Ultra accelerator achieves 75.4% and 3.5% savings in area and power, respectively, over regular SRAM-based accelerators.