End to End Learning for Self-Driving Cars

End to End Learning for Self-Driving Cars
复制标题

DOI:
--
复制
发表时间:
2016-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Mariusz Bojarski;D. Testa;Daniel Dworakowski;Bernhard Firner;B. Flepp;Prasoon Goyal;L. Jackel;
Mariusz Bojarski;D. Testa;Daniel Dworakowski;Bernhard Firner;B. Flepp;Prasoon Goyal;L. Jackel;
中科院分区:
其他
文献类型:
--
作者:
Mariusz Bojarski;D. Testa;Daniel Dworakowski;Bernhard Firner;B. Flepp;Prasoon Goyal;L. Jackel;

文献摘要

被引文献

相似文献

我们训练了一个卷积神经网络(CNN),将来自单个前置摄像头的原始像素直接映射到转向命令。事实证明,这种端到端的方法非常强大。通过最少的人类训练数据,该系统可以学习在有或没有车道标记的当地道路和高速公路上驾驶。它还可以在视觉引导不清晰的区域运行,例如停车场和未铺设的道路。该系统自动学习必要处理步骤的内部表示,例如仅使用人类转向角作为训练信号来检测有用的道路特征。我们从来没有明确地训练它来检测,例如,道路的轮廓。与问题的显式分解(如车道标记检测、路径规划和控制)相比,我们的端到端系统可同时优化所有处理步骤。我们认为,这将最终导致更好的性能和更小的系统。更好的性能将产生,因为内部组件自我优化以最大化整体系统性能,而不是优化人类选择的中间标准,例如,车道检测可以理解的是,选择这样的标准是为了便于人类解释,这并不能自动保证最大的系统性能。更小的网络是可能的,因为系统学习用最少的处理步骤来解决问题。我们使用NVIDIA DevBox和Torch 7进行训练,使用NVIDIA DRIVE(TM)PX自动驾驶汽车计算机也运行Torch 7来确定驾驶地点。该系统以每秒30帧(FPS)的速度运行。
We trained a convolutional neural network (CNN) to map raw pixels from a single front-facing camera directly to steering commands. This end-to-end approach proved surprisingly powerful. With minimum training data from humans the system learns to drive in traffic on local roads with or without lane markings and on highways. It also operates in areas with unclear visual guidance such as in parking lots and on unpaved roads. The system automatically learns internal representations of the necessary processing steps such as detecting useful road features with only the human steering angle as the training signal. We never explicitly trained it to detect, for example, the outline of roads. Compared to explicit decomposition of the problem, such as lane marking detection, path planning, and control, our end-to-end system optimizes all processing steps simultaneously. We argue that this will eventually lead to better performance and smaller systems. Better performance will result because the internal components self-optimize to maximize overall system performance, instead of optimizing human-selected intermediate criteria, e.g., lane detection. Such criteria understandably are selected for ease of human interpretation which doesn't automatically guarantee maximum system performance. Smaller networks are possible because the system learns to solve the problem with the minimal number of processing steps. We used an NVIDIA DevBox and Torch 7 for training and an NVIDIA DRIVE(TM) PX self-driving car computer also running Torch 7 for determining where to drive. The system operates at 30 frames per second (FPS).