Jointly Learning to Construct and Control Agents using Deep Reinforcement Learning

Jointly Learning to Construct and Control Agents using Deep Reinforcement Learning
复制标题

DOI:
10.1109/icra.2019.8793537
复制
发表时间:
2018-01
期刊:
2019 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Charles B. Schaff;David Yunis;Ayan Chakrabarti;Matthew R. Walter
Charles B. Schaff;David Yunis;Ayan Chakrabarti;Matthew R. Walter
中科院分区:
其他
文献类型:
--
作者:
Charles B. Schaff;David Yunis;Ayan Chakrabarti;Matthew R. Walter

文献摘要

相似文献

机器人的物理设计和控制其运动的策略是内在地耦合的,应该根据任务和环境来确定。在越来越多的应用中,数据驱动和基于学习的方法,如深度强化学习,在设计控制策略方面已被证明是有效的。对于大多数任务,根据此类控制策略评估物理设计的唯一方法是经验性的--即选择设计并为其培训控制策略。由于训练这些策略是耗时的,因此在计算上不可能为所有可能的设计训练单独的策略作为确定最佳设计的方法。在这项工作中,我们通过引入一种在物理设计和控制网络上联合优化的方法来解决这一限制。我们的方法保持了设计的分布,并使用强化学习来优化控制策略,以最大化设计分布的预期回报。我们允许控制器访问设计参数,以使其能够针对分发中的每个设计定制其策略。在整个培训过程中,我们将分布转向更高性能的设计,最终收敛到共同最优的设计和控制策略。我们在腿部运动的背景下评估了我们的方法,并证明了它发现了新颖的设计和行走步态,在不同的设置下表现优于基线。
The physical design of a robot and the policy that controls its motion are inherently coupled, and should be determined according to the task and environment. In an increasing number of applications, data-driven and learning-based approaches, such as deep reinforcement learning, have proven effective at designing control policies. For most tasks, the only way to evaluate a physical design with respect to such control policies is empirical—i.e., by picking a design and training a control policy for it. Since training these policies is time-consuming, it is computationally infeasible to train separate policies for all possible designs as a means to identify the best one. In this work, we address this limitation by introducing a method that jointly optimizes over the physical design and control network. Our approach maintains a distribution over designs and uses reinforcement learning to optimize a control policy to maximize expected reward over the design distribution. We give the controller access to design parameters to allow it to tailor its policy to each design in the distribution. Throughout training, we shift the distribution towards higher-performing designs, eventually converging to a design and control policy that are jointly optimal. We evaluate our approach in the context of legged locomotion, and demonstrate that it discovers novel designs and walking gaits, outperforming baselines across different settings.