Energy-Based Legged Robots Terrain Traversability Modeling via Deep Inverse Reinforcement Learning

Energy-Based Legged Robots Terrain Traversability Modeling via Deep Inverse Reinforcement Learning
复制标题

DOI:
10.1109/lra.2022.3188100
复制
发表时间:
2022-07
影响因子:
5.2
通讯作者:
Lu Gan;J. Grizzle;R. Eustice;Maani Ghaffari
Lu Gan;J. Grizzle;R. Eustice;Maani Ghaffari
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lu Gan;J. Grizzle;R. Eustice;Maani Ghaffari

文献摘要

被引文献

相似文献

这项工作报告了一种针对腿部机器人的深度强化学习方法,该方法既包含了扩展和本体感知的感觉数据,又使用机器人 - 敏感性的扩展环境特征或手工制作的运动功能;从一个深度神经网络中奖励近似的特征。改善模型保真度,并提供奖励,取决于机器人在部署期间的状态。示范。节能策略比示范使用了MIT MIN-CHEETAH机器人和迷你Cheetah模拟器收集的数据集。
This work reports ondeveloping a deep inverse reinforcement learning method for legged robots terrain traversability modeling that incorporates both exteroceptive and proprioceptive sensory data. Existing works use robot-agnostic exteroceptive environmental features or handcrafted kinematic features; instead, we propose to also learn robot-specific inertial features from proprioceptive sensory data for reward approximation in a single deep neural network. Incorporating the inertial features can improve the model fidelity and provide a reward that depends on the robot’s state during deployment. We train the reward network using the Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) algorithm and propose simultaneously minimizing a trajectory ranking loss to deal with the suboptimality of legged robot demonstrations. The demonstrated trajectories are ranked by locomotion energy consumption, in order to learn an energy-aware reward function and a more energy-efficient policy than demonstration. We evaluate our method using a dataset collected by an MIT Mini-Cheetah robot and a Mini-Cheetah simulator. The code is publicly available.1