Weighing in on photonic-based machine learning for automotive mobility

Weighing in on photonic-based machine learning for automotive mobility
复制标题

权衡基于光子的汽车移动机器学习

DOI:
10.1038/s41566-020-00736-0
复制
发表时间:
2021
期刊:
影响因子:
35
通讯作者:
E. Dede
E. Dede
中科院分区:
物理与天体物理1区
文献类型:
--
作者:
S. Rodrigues;Ziqi Yu;P. Schmalenberg;J. Lee;H. Iizuka;E. Dede

文献摘要

参考文献

被引文献

相似文献

人工神经网络的光学处理正在重新评估其解决互补金属氧化物半导体(CMOS)技术所面临的大功耗、数据穿梭和散热挑战的潜力。虽然最近的进展表明,每单位能量的光子计算量非常高,但光子处理必须在现代市场中找到一席之地。在需要有限的数值精度,高推理速度和并行性的应用中,这种替代计算框架可以优于传统的冯诺依曼架构。在数字化转型的转折点上,互联产品提供了对用户体验的洞察,光学处理可能是边缘计算与更高带宽的移动的网络相抗衡的关键。大约在1993年,当苹果公司售出第1000万台运行在400 kB软盘驱动器上的Macintosh时,实时面部识别技术出现在加州理工学院(Caltech)。虽然我们都熟悉今天在一些智能手机的微处理器上运行的面部识别软件,但加州理工学院的早期演示是建立在光子机器学习的基础上的。在人脸识别系统的基础上,人工神经网络识别特征以计算匹配图像的概率。虽然神经网络背后的理论是在20世纪40年代定义的,但第一台计算机尚未开发出来,算法的实用性也被推迟了。为了实现今天的面部识别,我们使用大型服务器来搜索数据库。相比之下,上述20世纪90年代的技术利用由光折射晶体组成的自由空间光学处理单元来存储训练的权重加上用于相关计算的透镜、反射镜和光栅。就在加州理工学院发表论文的一年前,马萨诸塞州理工学院也展示了使用代码和电子处理器的面部识别。用于机器学习的光子处理在很短的时间内就能与电子硬件相匹配。那么,为什么要挖掘35年前的技术呢?在当今的冯·诺依曼架构上,深度学习任务需要巨大的计算能力。考虑卷积神经网络在4 GPU系统上的训练过程; 1周的平均功率约为180 W。这相当于约32千瓦时的能源支出或分布式处理的更高需求。相比之下,根据美国能源信息署和欧盟统计局的数据,2018年美国和欧洲家庭平均每周消耗约211千瓦时和约21千瓦时。这些家庭价值相当于训练有素的机器学习工程师每周的耗电量。根据这一统计数据,IBM研究副总裁Mukesh Khare在2019年的一次神经形态会议上推测,到2040年,神经网络消耗的电力可能超过全球产生的总电力。如今,互联产品通过聚合数据利用对用户体验的洞察,有两种方法可以处理这些数据。第一种是将其发送到云计算服务,而第二种是在边缘处理它。未来的移动解决方案可能依赖于光探测和测距(LiDAR)技术,该技术可以快速生成大量数据(例如,2 TB min-1)。即使在移动的网络中取得了进步,处理这种连续的数据流也是笨拙的。此外,虽然边缘计算受到本地化功耗和硬件去本地化的限制,但它在隐私,连接和延迟方面具有优势。无论是在云端还是在路上,图形处理单元(GPU)都被认为是计算神经网络的标准硬件,尽管它是“通用”技术。相比之下,新的专用张量处理单元(TPU)在其基本计算单元处执行矩阵乘法。TPU通过较少的灵活性,降低了GPU的冯诺依曼处理速度瓶颈。然而,对于未来的汽车来说,最大化计算速度和最小化功耗至关重要。原始设备制造商丰田和通用汽车以及英伟达、恩智浦半导体和Arm等公司最近加入了一个自动驾驶汽车计算联盟,为未来的汽车设定要求。尽管有这些发展,下一代车队继续优先考虑处理器,暴露出快速,功率敏捷的边缘计算存在的技术差距。鉴于这一差距,光子处理器能否满足自主移动的需求?20世纪90年代的光神经网络的光子处理从未达到市场成熟,这是由于庞大的光学系统、较差的材料非线性以及不敏感的窄带宽光电处理。今天,硅光子学和电光转换的发展带来了模拟计算光子学的复兴。最近的进展已经证明了多种方法可以加速光子处理以实现定制的计算应用。这些包括光学尖峰神经形态系统3,实现为无损系统的集成电路,线性光学operators 4 -6,用于广播和重量架构的波分复用7,以及自由空间衍射元表面8。所有这些技术看起来云计算边缘计算
To the Editor — Optical processing for artificial neural networks is being re-evaluated for its potential to address large power consumption, data shuttling and thermal dissipation challenges faced by complementary metal–oxide– semiconductor (CMOS) technology. Though recent advances demonstrate an exceptionally high number of photonic calculations per unit energy, photonic processing must find a niche in the modern market. In applications requiring limited numeric precision, high inference speed and parallelism, this alternative computing framework could outperform traditional von Neumann architectures. At the inflection point of a digital transformation where connected products provide insight into the user experience, optical processing may be the key to pit edge computing against higher-bandwidth mobile networks. Around 1993 when Apple sold its ten millionth Macintosh that ran on a 400 kB floppy drive, real-time facial recognition emerged at the California Institute of Technology (Caltech)1. While we are all familiar with today’s facial recognition software that runs on microprocessors in some smartphones, this early demonstration by Caltech was built on a form of photonic machine learning. At the base of facial recognition systems, an artificial neural network identifies features to calculate the probability of a matching image. While the theory behind neural networks was defined in the 1940s, the first computers were not yet developed, and the utility of the algorithm was delayed. To achieve facial recognition today, we enlist large servers to scour databases. In comparison, the aforementioned 1990s technology utilized free-space optical processing units consisting of photorefractive crystals to store the trained weights plus lenses, mirrors and gratings for associated calculations. Just a year before Caltech’s publication, the Massachusetts Institute of Technology also demonstrated facial recognition using code and an electronic processor. Photonic processing for machine learning had matched electronic hardware in little time. So why unearth technology from 35 years ago? Tremendous computational power is needed for deep-learning tasks on today’s von Neumann architecture. Consider the training process of a convolutional neural network on a 4 GPU system; an average power for 1 week is about 180 W. This amounts to ~32 kWh of energy expenditure or higher demands for distributed processing. Comparatively, the average American and European household in 2018 consumed ~211 kWh per week and ~21 kWh per week, according to the US Energy Information Administration and Eurostat. These household values are on par with the weekly power consumption of a trained machine-learning engineer. Extrapolating this statistic, IBM vice president of research, Mukesh Khare, speculated at a neuromorphic conference in 2019 that power consumed by neural networks could exceed the total power generated in the world by 20402. Today, connected products leverage insight into the user experience via aggregated data and there are two methods to work with this volume of data. The first is to send it to a cloud computing service, while the second is to process it at the edge. Future mobility solutions are likely to rely on light detection and ranging (LiDAR) technology, which is known to rapidly generate massive amounts of data (for example, 2 TB min–1). Even with advancements in mobile networks, handling this continuous stream of data is unwieldy. Furthermore, while edge computing is limited by localized power consumption and hardware delocalization, it has advantages in privacy, connectivity and latency. Whether in the cloud or on the road, graphic processing units (GPUs) have been considered the standard hardware to calculate neural networks despite being ‘general use’ technology. In contrast, new specialized tensor processing units (TPUs) perform matrix multiplication at their base unit of calculation. The TPU, through less flexibility, reduces the von Neumann processing-speed bottleneck of a GPU. However, for future vehicles, maximizing computing speed and minimizing power consumption is critical. Original equipment manufacturers, Toyota and General Motors, and firms such as NVIDIA, NXP Semiconductors and Arm recently joined an Autonomous Vehicle Computing Consortium to set requirements for future vehicles. Despite these developments, the next fleet of vehicles continues to prioritize power over processor, baring a technology gap where fast, power-agile edge computing resides. In light of this gap, can a photonic processor address the needs of autonomous mobility? Photonic processing of optical neural networks in the 1990s never reached market maturity due to bulky optical systems, poor material nonlinearity, and insensitive, narrow-bandwidth optoelectronic processing. Today, developments in silicon photonics and electro-optic conversion have brought about a resurgence in photonics for analogue computing. Recent advances have demonstrated a variety of methods to accelerate photonic processing for tailored computing applications. These include optical spiking neuromorphic systems3, integrated circuits implemented as a system of lossless, linear optical operators4–6, wavelength-division multiplexing for broadcast and weight architectures7, and free-space diffractive metasurfaces8. All these technologies look Cloud computing Edge computing
DOI: 10.1038/nphoton.2017.93
发表时间: 2017-07-01
期刊: NATURE PHOTONICS
影响因子: 35
作者:
Shen, Yichen;Harris, Nicholas C.;Soljacic, Marin
通讯作者: Soljacic, Marin