Single chip photonic deep neural network with accelerated training

Single chip photonic deep neural network with accelerated training
复制标题

DOI:
10.48550/arxiv.2208.01623
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Bandyopadhyay;Alexander Sludds;Stefan Krastanov;R. Hamerly;N. Harris;D. Bunandar;M. Streshinsky;M. Hochberg;D. Englund
S. Bandyopadhyay;Alexander Sludds;Stefan Krastanov;R. Hamerly;N. Harris;D. Bunandar;M. Streshinsky;M. Hochberg;D. Englund
中科院分区:
其他
文献类型:
--
作者:
S. Bandyopadhyay;Alexander Sludds;Stefan Krastanov;R. Hamerly;N. Harris;D. Bunandar;M. Streshinsky;M. Hochberg;D. Englund

文献摘要

被引文献

相似文献

随着深度神经网络(dnn)彻底改变机器学习,能耗和吞吐量正在成为CMOS电子产品的基本限制。这促使人们寻找针对人工智能优化的新硬件架构,如电子收缩阵列、忆阻器交叉栅阵列和光学加速器。光学系统可以以极高的速率和效率执行线性矩阵运算,这激发了最近低延迟线性代数和每次乘法累积操作低于一个光子的光能消耗的演示。然而,在单个芯片中演示线性和非线性处理单元的协集成系统仍然是一个主要挑战。在这里,我们在可扩展光子集成电路(PIC)中引入了这样一个系统,通过几个关键进展:(i)高带宽和低功耗可编程非线性光学功能单元(NOFUs);(ii)相干矩阵乘法单元(CMXUs);(三)利用光学加速度进行现场训练。我们通过实验证明了这种完全集成的相干光神经网络(FICONN)架构适用于一个三层DNN,该DNN由12个nofu和3个cmxu组成,运行在电信c波段。通过对元音分类任务的原位训练,FICONN在测试集上达到了92.7%的准确率,这与具有相同权重数的数字计算机上获得的准确率相同。这项工作为原位训练的理论建议提供了实验证据,解锁了训练数据吞吐量的数量级改进。此外,FICONN开启了以纳秒延迟和飞焦耳每次操作能量效率进行推断的道路。
As deep neural networks (DNNs) revolutionize machine learning, energy consumption and throughput are emerging as fundamental limitations of CMOS electronics. This has motivated a search for new hardware architectures optimized for artificial intelligence, such as electronic systolic arrays, memristor crossbar arrays, and optical accelerators. Optical systems can perform linear matrix operations at exceptionally high rate and efficiency, motivating recent demonstrations of low latency linear algebra and optical energy consumption below a photon per multiply-accumulate operation. However, demonstrating systems that co-integrate both linear and nonlinear processing units in a single chip remains a central challenge. Here we introduce such a system in a scalable photonic integrated circuit (PIC), enabled by several key advances: (i) high-bandwidth and low-power programmable nonlinear optical function units (NOFUs); (ii) coherent matrix multiplication units (CMXUs); and (iii) in situ training with optical acceleration. We experimentally demonstrate this fully-integrated coherent optical neural network (FICONN) architecture for a 3-layer DNN comprising 12 NOFUs and three CMXUs operating in the telecom C-band. Using in situ training on a vowel classification task, the FICONN achieves 92.7% accuracy on a test set, which is identical to the accuracy obtained on a digital computer with the same number of weights. This work lends experimental evidence to theoretical proposals for in situ training, unlocking orders of magnitude improvements in the throughput of training data. Moreover, the FICONN opens the path to inference at nanosecond latency and femtojoule per operation energy efficiency.