TACO: Learning Task Decomposition via Temporal Alignment for Control

TACO: Learning Task Decomposition via Temporal Alignment for Control
复制标题

TACO:通过时间对齐进行控制的学习任务分解

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
I. Posner
I. Posner
中科院分区:
--
文献类型:
--
作者:
K. Shiarlis;Markus Wulfmeier;Sasha Salter;Shimon Whiteson;I. Posner

文献摘要

被引文献

相似文献

许多先进的演示学习 (LfD) 方法考虑将复杂的现实世界任务分解为更简单的子任务。通过在任务内和任务之间重用相应的子策略,它们为来自不同高级任务的每个策略提供训练数据,并将它们组合起来执行新颖的任务。现有的模块化 LfD 方法要么专注于学习单个高级任务,要么依赖于领域知识和时间分割。相比之下,我们提出了一种基于任务草图的弱监督、与领域无关的方法,其中仅包括每个演示中执行的子任务的序列。我们的方法同时将草图与观察到的演示进行对齐,并学习所需的子策略。与单独的优化过程相比,这提高了泛化能力。我们在多个领域评估该方法,包括使用纯粹基于图像的观察来模拟 3D 机器人手臂控制任务。结果表明,我们的方法与完全监督的方法表现相当,同时需要的注释工作显着减少。
Many advanced Learning from Demonstration (LfD) methods consider the decomposition of complex, real-world tasks into simpler sub-tasks. By reusing the corresponding sub-policies within and between tasks, they provide training data for each policy from different high-level tasks and compose them to perform novel ones. Existing approaches to modular LfD focus either on learning a single high-level task or depend on domain knowledge and temporal segmentation. In contrast, we propose a weakly supervised, domain-agnostic approach based on task sketches, which include only the sequence of sub-tasks performed in each demonstration. Our approach simultaneously aligns the sketches with the observed demonstrations and learns the required sub-policies. This improves generalisation in comparison to separate optimisation procedures. We evaluate the approach on multiple domains, including a simulated 3D robot arm control task using purely image-based observations. The results show that our approach performs commensurately with fully supervised approaches, while requiring significantly less annotation effort.