Toward a Behavioral-Level End-to-End Framework for Silicon Photonics Accelerators

Toward a Behavioral-Level End-to-End Framework for Silicon Photonics Accelerators
复制标题

DOI:
10.1109/igsc55832.2022.9969371
复制
发表时间:
2022-10
期刊:
2022 IEEE 13th International Green and Sustainable Computing Conference (IGSC)
影响因子:
--
通讯作者:
Emily Lattanzio;Ranyang Zhou;A. Roohi;Abdallah Khreishah;Durga Misra;Shaahin Angizi
Emily Lattanzio;Ranyang Zhou;A. Roohi;Abdallah Khreishah;Durga Misra;Shaahin Angizi
中科院分区:
其他
文献类型:
--
作者:
Emily Lattanzio;Ranyang Zhou;A. Roohi;Abdallah Khreishah;Durga Misra;Shaahin Angizi

文献摘要

相似文献

卷积神经网络(CNN)由于其在各种AI应用(例如对象识别、语音处理等)中的有效性而被广泛使用,其中乘法和累加(MAC)操作贡献了$\sim 95\%$的计算时间。从硬件实现的角度来看,目前基于CMOS的MAC加速器的性能是有限的,主要是由于其冯诺依曼架构和相应的有限的内存带宽。通过这种方式,硅光子学最近被探索为加速器设计的有前途的解决方案,以提高设计的速度和功率效率,而不是电子忆阻交叉开关。在这项工作中,我们简要地研究了最近的硅光子加速器,并采取初步措施,开发一个开源和自适应交叉结构模拟器。在保留MNSIM工具[1]的原始功能的基础上,我们添加了一种新的光子模式,该模式利用预先存在的算法与基于光子相变存储器(pPCM)的交叉开关结构一起工作。通过CNN的拓扑结构、加速器配置和实验基准数据的输入,所呈现的模拟器可以报告最佳交叉开关大小、所需交叉开关的数量以及总面积、功率和延迟的估计。
Convolutional Neural Networks (CNNs) are widely used due to their effectiveness in various AI applications such as object recognition, speech processing, etc., where the multiply-and-accumulate (MAC) operation contributes to $\sim 95\%$ of the computation time. From the hardware implementation perspective, the performance of current CMOS-based MAC accelerators is limited mainly due to their von-Neumann architecture and corresponding limited memory bandwidth. In this way, silicon photonics has been recently explored as a promising solution for accelerator design to improve the speed and power-efficiency of the designs as opposed to electronic memristive crossbars. In this work, we briefly study recent silicon photonics accelerators and take initial steps to develop an open-source and adaptive crossbar architecture simulator for that. Keeping the original functionality of the MNSIM tool [1], we add a new photonic mode that utilizes the pre-existing algorithm to work with a photonic Phase Change Memory (pPCM) based crossbar structure. With inputs from the CNN's topology, the accelerator configuration, and experimentally-benchmarked data, the presented simulator can report the optimal crossbar size, the number of crossbars needed, and the estimation of total area, power, and latency.