Benchmarking Structured Policies and Policy Optimization for Real-World Dexterous Object Manipulation

Benchmarking Structured Policies and Policy Optimization for Real-World Dexterous Object Manipulation
复制标题

DOI:
10.1109/lra.2021.3129139
复制
发表时间:
2021-05
影响因子:
5.2
通讯作者:
Niklas Funk;Charles B. Schaff;Rishabh Madan;Takuma Yoneda;Julen Urain De Jesus;Joe Watson;E. Gordon;F. Widmaier;Stefan Bauer;S. Srinivasa;T. Bhattacharjee;Matthew R. Walter;Jan Peters
Niklas Funk;Charles B. Schaff;Rishabh Madan;Takuma Yoneda;Julen Urain De Jesus;Joe Watson;E. Gordon;F. Widmaier;Stefan Bauer;S. Srinivasa;T. Bhattacharjee;Matthew R. Walter;Jan Peters
中科院分区:
计算机科学2区
文献类型:
--
作者:
Niklas Funk;Charles B. Schaff;Rishabh Madan;Takuma Yoneda;Julen Urain De Jesus;Joe Watson;E. Gordon;F. Widmaier;Stefan Bauer;S. Srinivasa;T. Bhattacharjee;Matthew R. Walter;Jan Peters

文献摘要

相似文献

灵巧操作是机器人技术中一个具有挑战性的重要问题。虽然数据驱动方法是一种很有前途的方法,但由于流行方法的样本效率低下,目前的基准测试需要模拟或广泛的工程支持。我们为TriFinger系统提供了基准测试,TriFinger系统是一个用于灵巧操作的开源机器人平台,也是2020年真实的机器人挑战赛的焦点。在挑战中取得成功的基准方法通常可以被描述为结构化策略,因为它们结合了经典机器人技术和现代策略优化的联合收割机元素。这种包括电感偏置有利于样品效率,可解释性,可靠性和高性能。该基准测试的关键方面是验证模拟和真实的系统的基线,对每个解决方案的核心功能进行彻底的消融研究,以及对作为操作基准的挑战进行回顾性分析。
Dexterous manipulation is a challenging and important problem in robotics. While data-driven methods are a promising approach, current benchmarks require simulation or extensive engineering support due to the sample inefficiency of popular methods. We present benchmarks for the TriFinger system, an open-source robotic platform for dexterous manipulation and the focus of the 2020 Real Robot Challenge. The benchmarked methods, which were successful in the challenge, can be generally described as structured policies, as they combine elements of classical robotics and modern policy optimization. This inclusion of inductive biases facilitates sample efficiency, interpretability, reliability and high performance. The key aspects of this benchmarking is validation of the baselines across both simulation and the real system, thorough ablation study over the core features of each solution, and a retrospective analysis of the challenge as a manipulation benchmark.