Air Learning: An AI Research Platform for Algorithm-Hardware Benchmarking of Autonomous Aerial Robots

Air Learning: An AI Research Platform for Algorithm-Hardware Benchmarking of Autonomous Aerial Robots
复制标题

Air Learning:自主空中机器人算法硬件基准测试的人工智能研究平台

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
V. Reddi
V. Reddi
中科院分区:
--
文献类型:
--
作者:
Srivatsan Krishnan;Behzad Boroujerdian;William Fu;Aleksandra Faust;V. Reddi

文献摘要

被引文献

相似文献

我们介绍了空气学习,这是一个人工智能研究平台,用于对算法硬件性能和能效权衡进行基准测试。我们特别关注自主无人驾驶飞行器(uav)中的深度强化学习(RL)交互。配备随机环境生成器,AirLearning将无人机暴露在各种具有挑战性的场景中。用户可以指定任务,训练不同的RL策略,并在各种硬件平台上评估其性能和能效。为了展示如何使用空气学习,我们将其与深度Q网络(DQN)和近端策略优化(PPO)一起播种,以解决使用我们的可配置环境生成器生成的三种不同环境中的点对点避障任务。我们使用课程学习和非课程学习来训练这两种算法。Air Learning在资源受限的嵌入式平台(如Ras-Pi)上,根据各种飞行质量(QoF)指标(如能耗、续航时间和平均轨迹长度)评估经过训练的策略的性能。我们发现嵌入式Ras-Pi上的轨迹与高端桌面系统上的轨迹有很大不同,导致其中一种环境下的轨迹最长可达79.43%。为了理解这种差异的根源,我们使用Air Learning人为地降低桌面性能,以模拟低端嵌入式系统上发生的情况。硬件在环的QoF指标描述了这些差异,并揭示了机载计算的选择如何影响空中机器人的性能。我们还进行了可靠性研究,以证明空气学习如何帮助理解传感器故障如何影响学习策略。所有这些放在一起,空中学习使无人机的RL研究成为可能。更多关于空气学习的信息和代码可以在这里找到:这个http URL
We introduce Air Learning, an AI research platform for benchmarking algorithm-hardware performance and energy efficiency trade-offs. We focus in particular on deep reinforcement learning (RL) interactions in autonomous unmanned aerial vehicles (UAVs). Equipped with a random environment generator, AirLearning exposes a UAV to a diverse set of challenging scenarios. Users can specify a task, train different RL policies and evaluate their performance and energy efficiency on a variety of hardware platforms. To show how Air Learning can be used, we seed it with Deep Q Networks (DQN) and Proximal Policy Optimization (PPO) to solve a point-to-point obstacle avoidance task in three different environments, generated using our configurable environment generator. We train the two algorithms using curriculum learning and non-curriculum-learning. Air Learning assesses the trained policies' performance, under a variety of quality-of-flight (QoF) metrics, such as the energy consumed, endurance and the average trajectory length, on resource-constrained embedded platforms like a Ras-Pi. We find that the trajectories on an embedded Ras-Pi are vastly different from those predicted on a high-end desktop system, resulting in up to 79.43% longer trajectories in one of the environments. To understand the source of such differences, we use Air Learning to artificially degrade desktop performance to mimic what happens on a low-end embedded system. QoF metrics with hardware-in-the-loop characterize those differences and expose how the choice of onboard compute affects the aerial robot's performance. We also conduct reliability studies to demonstrate how Air Learning can help understand how sensor failures affect the learned policies. All put together, Air Learning enables a broad class of RL studies on UAVs. More information and code for Air Learning can be found here: this http URL