Real-time multi-task diffractive deep neural networks via hardware-software co-design.

Real-time multi-task diffractive deep neural networks via hardware-software co-design.
复制标题

DOI:
10.1038/s41598-021-90221-7
复制
发表时间:
2021-05-26
期刊:
影响因子:
4.6
通讯作者:
Yu C
Yu C
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Li Y;Chen R;Sensale-Rodriguez B;Gao W;Yu C

文献摘要

被引文献

相似文献

深度神经网络(DNNS)具有实质性的计算要求,这极大地限制了其在资源受限环境中的性能。最近,在光学神经网络和基于光学计算的DNN硬件方面正在越来越多的努力,这在其功率效率,并行性和计算速度方面为深度学习系统带来了重要优势。其中,基于光衍射的自由空间衍射深神经网络(D2NN),在相邻层中与神经元相互关联的每一层中具有数百万个神经元。但是,由于实施可重新配置的挑战,部署不同的DNN算法需要重新构建并复制物理衍射系统,这大大降低了实际应用程序场景中的硬件效率。因此,这项工作提出了一种新颖的硬件 - 软件共同设计方法,该方法可以在D22NN中实现类似其首先的实时多任务学习,从而自动识别实时部署哪些任务。我们的实验结果表明,在所有系统组件的宽噪声范围内,多功能性,硬件效率的显着提高,并证明和量化了建议的多任务D2NN体系结构的鲁棒性。此外,我们为训练提出的多任务架构提供了一种特定领域的正则化算法,该算法可灵活地调整每个任务的所需性能。
Deep neural networks (DNNs) have substantial computational requirements, which greatly limit their performance in resource-constrained environments. Recently, there are increasing efforts on optical neural networks and optical computing based DNNs hardware, which bring significant advantages for deep learning systems in terms of their power efficiency, parallelism and computational speed. Among them, free-space diffractive deep neural networks (D2NNs) based on the light diffraction, feature millions of neurons in each layer interconnected with neurons in neighboring layers. However, due to the challenge of implementing reconfigurability, deploying different DNNs algorithms requires re-building and duplicating the physical diffractive systems, which significantly degrades the hardware efficiency in practical application scenarios. Thus, this work proposes a novel hardware-software co-design method that enables first-of-its-like real-time multi-task learning in D22NNs that automatically recognizes which task is being deployed in real-time. Our experimental results demonstrate significant improvements in versatility, hardware efficiency, and also demonstrate and quantify the robustness of proposed multi-task D2NN architecture under wide noise ranges of all system components. In addition, we propose a domain-specific regularization algorithm for training the proposed multi-task architecture, which can be used to flexibly adjust the desired performance for each task.