Learning strategies in table tennis using inverse reinforcement learning

Learning strategies in table tennis using inverse reinforcement learning
复制标题

DOI:
10.1007/s00422-014-0599-1
复制
发表时间:
2014-10-01
影响因子:
1.9
通讯作者:
Peters, Jan
Peters, Jan
中科院分区:
工程技术3区
文献类型:
--
作者:
Muelling, Katharina;Boularias, Abdeslam;Peters, Jan

文献摘要

被引文献

相似文献

对于机器人和人类来说,学习乒乓球这样的复杂任务都是一个具有挑战性的问题。即使在获得了必要的运动技能之后,也需要一个策略来选择在哪里以及如何将球送到对手的球场上,才能赢得比赛。在乒乓球等互动任务中,通过数据驱动的基本策略识别在很大程度上是一个未被探索的问题。在本文中,我们提出了一个基于马尔可夫决策问题的策略表示和推理的计算模型,其中奖励函数对任务的目标和策略信息进行建模。我们展示了如何使用无模型的逆强化学习从乒乓球比赛的演示中发现这种奖励函数。由此产生的框架允许确定选择击打动作所依据的基本要素。我们从不同比赛风格和不同比赛条件下收集的数据上测试了我们的方法。估计奖励功能能够捕获特定于专家的战略信息,这些信息足以在不同技能水平和不同打法风格的球员中区分专家。
Learning a complex task such as table tennis is a challenging problem for both robots and humans. Even after acquiring the necessary motor skills, a strategy is needed to choose where and how to return the ball to the opponent's court in order to win the game. The data-driven identification of basic strategies in interactive tasks, such as table tennis, is a largely unexplored problem. In this paper, we suggest a computational model for representing and inferring strategies, based on a Markov decision problem, where the reward function models the goal of the task as well as the strategic information. We show how this reward function can be discovered from demonstrations of table tennis matches using model-free inverse reinforcement learning. The resulting framework allows to identify basic elements on which the selection of striking movements is based. We tested our approach on data collected from players with different playing styles and under different playing conditions. The estimated reward function was able to capture expert-specific strategic information that sufficed to distinguish the expert among players with different skill levels as well as different playing styles.