Weighing in on photonic-based machine learning for automotive mobility
Weighing in on photonic-based machine learning for automotive mobility
复制标题
权衡基于光子的汽车移动机器学习
DOI:
10.1038/s41566-020-00736-0
复制
发表时间:
2021
期刊:
影响因子:
35
通讯作者:
E. Dede
中科院分区:
文献类型:
--
作者:
S. Rodrigues;Ziqi Yu;P. Schmalenberg;J. Lee;H. Iizuka;E. Dede
To the Editor — Optical processing for artificial neural networks is being re-evaluated for its potential to address large power consumption, data shuttling and thermal dissipation challenges faced by complementary metal–oxide– semiconductor (CMOS) technology. Though recent advances demonstrate an exceptionally high number of photonic calculations per unit energy, photonic processing must find a niche in the modern market. In applications requiring limited numeric precision, high inference speed and parallelism, this alternative computing framework could outperform traditional von Neumann architectures. At the inflection point of a digital transformation where connected products provide insight into the user experience, optical processing may be the key to pit edge computing against higher-bandwidth mobile networks. Around 1993 when Apple sold its ten millionth Macintosh that ran on a 400 kB floppy drive, real-time facial recognition emerged at the California Institute of Technology (Caltech)1. While we are all familiar with today’s facial recognition software that runs on microprocessors in some smartphones, this early demonstration by Caltech was built on a form of photonic machine learning. At the base of facial recognition systems, an artificial neural network identifies features to calculate the probability of a matching image. While the theory behind neural networks was defined in the 1940s, the first computers were not yet developed, and the utility of the algorithm was delayed. To achieve facial recognition today, we enlist large servers to scour databases. In comparison, the aforementioned 1990s technology utilized free-space optical processing units consisting of photorefractive crystals to store the trained weights plus lenses, mirrors and gratings for associated calculations. Just a year before Caltech’s publication, the Massachusetts Institute of Technology also demonstrated facial recognition using code and an electronic processor. Photonic processing for machine learning had matched electronic hardware in little time. So why unearth technology from 35 years ago? Tremendous computational power is needed for deep-learning tasks on today’s von Neumann architecture. Consider the training process of a convolutional neural network on a 4 GPU system; an average power for 1 week is about 180 W. This amounts to ~32 kWh of energy expenditure or higher demands for distributed processing. Comparatively, the average American and European household in 2018 consumed ~211 kWh per week and ~21 kWh per week, according to the US Energy Information Administration and Eurostat. These household values are on par with the weekly power consumption of a trained machine-learning engineer. Extrapolating this statistic, IBM vice president of research, Mukesh Khare, speculated at a neuromorphic conference in 2019 that power consumed by neural networks could exceed the total power generated in the world by 20402. Today, connected products leverage insight into the user experience via aggregated data and there are two methods to work with this volume of data. The first is to send it to a cloud computing service, while the second is to process it at the edge. Future mobility solutions are likely to rely on light detection and ranging (LiDAR) technology, which is known to rapidly generate massive amounts of data (for example, 2 TB min–1). Even with advancements in mobile networks, handling this continuous stream of data is unwieldy. Furthermore, while edge computing is limited by localized power consumption and hardware delocalization, it has advantages in privacy, connectivity and latency. Whether in the cloud or on the road, graphic processing units (GPUs) have been considered the standard hardware to calculate neural networks despite being ‘general use’ technology. In contrast, new specialized tensor processing units (TPUs) perform matrix multiplication at their base unit of calculation. The TPU, through less flexibility, reduces the von Neumann processing-speed bottleneck of a GPU. However, for future vehicles, maximizing computing speed and minimizing power consumption is critical. Original equipment manufacturers, Toyota and General Motors, and firms such as NVIDIA, NXP Semiconductors and Arm recently joined an Autonomous Vehicle Computing Consortium to set requirements for future vehicles. Despite these developments, the next fleet of vehicles continues to prioritize power over processor, baring a technology gap where fast, power-agile edge computing resides. In light of this gap, can a photonic processor address the needs of autonomous mobility? Photonic processing of optical neural networks in the 1990s never reached market maturity due to bulky optical systems, poor material nonlinearity, and insensitive, narrow-bandwidth optoelectronic processing. Today, developments in silicon photonics and electro-optic conversion have brought about a resurgence in photonics for analogue computing. Recent advances have demonstrated a variety of methods to accelerate photonic processing for tailored computing applications. These include optical spiking neuromorphic systems3, integrated circuits implemented as a system of lossless, linear optical operators4–6, wavelength-division multiplexing for broadcast and weight architectures7, and free-space diffractive metasurfaces8. All these technologies look Cloud computing Edge computing
影响因子:
35
作者:
Shen, Yichen;Harris, Nicholas C.;Soljacic, Marin
通讯作者:
Soljacic, Marin