Goal Babbling: a New Concept for Early Sensorimotor Exploration

Goal Babbling: a New Concept for Early Sensorimotor Exploration
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Matthias Rolf;Jochen J. Steil
Matthias Rolf;Jochen J. Steil
中科院分区:
其他
文献类型:
--
作者:
Matthias Rolf;Jochen J. Steil

文献摘要

被引文献

相似文献

人体拥有600多块骨骼肌[1]。为了实现某种行为目标而采取有目的的行动,需要高度协调这些多个自由度。然而,人类婴儿生来就没有最基本的协调技能,比如伸手拿东西[2],这使得感觉运动协调的学习成为人类发展中的一个基本问题。理解这种学习能力,并将其用于现代机器人系统是认知机器人和发展机器人研究领域的主要目标之一[4]、[5]。我们调查了作为一种典型的协调技能的到达技能的学习情况。伸展的问题是找到将手或机器人的末端执行器移动到空间中某一所需位置的电机命令(例如,机器人手臂的关节角度)。因此,运动指令q和结果x通过一种因果关系联系在一起,这被表示为正向函数f(Q)=x。学习需要颠倒这种关系以获得某些期望的结果x∗。这个问题设置不仅是说明性的,而且对于其他协调问题也是非常典型的:它提出了如何通过行动实现一些行为目标的非常一般性的问题。伸手的技能对机器人和人类来说都是基本的,因为在太空中的定位对于机器人的抓手或人类的手的任何使用都是必要的。成功的达成技能可以通过内部模型[6]、[7]的概念得到很好的理解,而正向模型预测行动的结果,反向模型建议行动以实现预期的结果。在没有明确的先验知识的情况下,内部模型的自举需要通过探索产生的经验。因此,机器学习方法传统上依赖于对所有可能的运动命令的详尽探索,这些命令通常是通过整个随机过程产生的,称为“运动喋喋不休”[8]、[9]。在数据生成阶段之后,可以用各种方式来表述学习和协调[10]、[11]、[12]。然而,在人体、现代类人机器人或象鼻等仿生机器人等高维运动系统上,不可能实现详尽的探索。用于不同执行器的命令组合的绝对数量太大,在任何学习代理的生命周期中都无法探索。了解人类运动的发展,以及未来机器人系统的成功应用,如仿生处理助手(见图3),需要成功的概念和方法
The human body possesses more than 600 skeletal muscles [1]. Performing purposeful actions to achieve some behavioral goal requires a high degree of coordination of these many degrees of freedom. Yet, human infants are born without the most basic coordination skills like reaching for an object [2], which poses the learning of sensorimotor coordination as a fundamental problem in human development. Understanding this ability to learn, and utilizing it for modern robotics systems is one of the major goals of the research fields of cognitive [3] and developmental robotics [4], [5]. We investigate the learning of reaching skills as an exemplary coordination skill. The problem of reaching is to find motor commands (e.g. joint angles of a robot arm) that move the hand, or the robot’s end-effector towards some desired position in space. Thereby motor commands q and outcomes x are connected by a causal relation which is denoted as the forward function f(q) = x. Learning needs to invert this relation in order achieve some desired outcome x∗. This problem setup is not only illustrative, but very prototypical for other coordination problems: it asks the very general question of how to achieve some behavioral goals by means of actions. The skill of reaching itself is also fundamental for both robots and humans, since the positioning in space is necessary for any use of the robot’s gripper or the human’s hand. Successful reaching skills can be well understood with the notion of internal models [6], [7], whereas forward models predict the outcome of an action and inverse models suggest actions in order to achieve a desired outcome. The bootstrapping of internal models without explicit prior-knowledge requires experience that has to be generated by exploration. Machine learning approaches thereby traditionally rely on an exhaustive exploration of all possible motor commands, frequently generated by means of an entire random procedure, which is referred to as “motor babbling” [8], [9]. After the data generation phase, learning and coordination can be phrased in a variety of ways [10], [11], [12]. Yet, exhaustive exploration can not be achieved on high-dimensional motor systems such as the human body, modern humanoid robots, or biomimetic robots like elephant trunks. The sheer number of combinations of commands for different actuators is too large to be explored in the lifetime of any learning agent. Understanding human motor development, as well as the successful application of future robotic systems like the Bionic Handling Assistant (see Fig. 3), demands for concepts and methods that succeed in