Learning Operational Space Control

Learning Operational Space Control
复制标题

学习操作空间控制

DOI:
--
复制
发表时间:
2006
期刊:
Robotics: Science and Systems
影响因子:
--
通讯作者:
S. Schaal
S. Schaal
中科院分区:
--
文献类型:
--
作者:
Jan Peters;S. Schaal

文献摘要

被引文献

相似文献

虽然操作空间控制对机器人技术至关重要,并且从分析的角度很好地理解,但面对建模误差,实现精确控制可能非常困难,这在复杂的机器人中是不可避免的,例如人形机器人。在这种情况下,学习控制方法可以提供一个有趣的替代分析控制算法。然而,由此产生的学习问题定义不清,因为它需要学习一个通常冗余系统的逆映射,众所周知,该系统具有解空间的非凸性,即学习系统可以生成电机命令,试图引导机器人进入物理上不可能的配置。本文的第一个重要见解是,尽管如此,当逆映射以适当的分段线性方式进行学习时,确实存在逆问题的物理正确解。我们工作的第二个关键组成部分是基于最近的一个见解,即许多操作空间控制器可以用约束最优控制问题来理解。与此最优控制问题相关的成本函数使我们能够制定一种学习算法,该算法在学习操作空间控制器的同时自动合成全局一致的期望冗余分辨率。从机器学习的角度来看,学习问题对应于一个强化学习问题,它最大化了即时奖励,并采用了期望最大化策略搜索算法。对一个三自由度机械臂的评估说明了所建议方法的可行性。
While operational space control is of essential importance for robotics and well-understood from an analytical point of view, it can be prohibitively hard to achieve accurate control in face of modeling errors, which are inevitable in complex robots, e.g., humanoid robots. In such cases, learning control methods can offer an interesting alternative to analytical control algorithms. However, the resulting learning problem is ill-defined as it requires to learn an inverse mapping of a usually redundant system, which is well known to suffer from the property of non-convexity of the solution space, i.e., the learning system could generate motor commands that try to steer the robot into physically impossible configurations. A first important insight for this paper is that, nevertheless, a physically correct solution to the inverse problem does exit when learning of the inverse map is performed in a suitable piecewise linear way. The second crucial component for our work is based on a recent insight that many operational space controllers can be understood in terms of a constraint optimal control problem. The cost function associated with this optimal control problem allows us to formulate a learning algorithm that automatically synthesizes a globally consistent desired resolution of redundancy while learning the operational space controller. From the view of machine learning, the learning problem corresponds to a reinforcement learning problem that maximizes an immediate reward and that employs an expectation-maximization policy search algorithm. Evaluations on a three degrees of freedom robot arm illustrate the feasibility of the suggested approach.