Computer Vision - ACCV 2020 - 15th Asian Conference on Computer Vision, Kyoto, Japan, November 30 - December 4, 2020, Revised Selected Papers, Part III

Computer Vision - ACCV 2020 - 15th Asian Conference on Computer Vision, Kyoto, Japan, November 30 - December 4, 2020, Revised Selected Papers, Part III
复制标题

计算机视觉 - ACCV 2020 - 第 15 届亚洲计算机视觉会议,日本京都,2020 年 11 月 30 日至 12 月 4 日,修订后的精选论文,第三部分

DOI:
10.1007/978-3-030-69535-4_44
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Nazarczuk M
Nazarczuk M
中科院分区:
--
文献类型:
--
作者:
Nazarczuk M

文献摘要

相似文献

在这项工作中,我们提出了一个新的人工智能任务——视觉到行动(V2A)——其中要求代理(机械臂)对场景中存在的对象(例如堆叠)执行高级任务。代理必须提出一个由原始动作(例如简单的移动、抓取)组成的计划,以便成功完成给定的任务。指令的制定方式迫使智能体在推断动作之前对所呈现的场景进行视觉推理。我们扩展了最近引入的数据集 SHOP-VRB,其中包含每个场景的任务指令以及能够评估基元序列是否导致任务成功完成的引擎。我们还针对此任务提出了一种基于多模态注意力的新方法,并在新数据集上展示了其性能。
In this work, we present a new AI task-Vision to Action (V2A)-where an agent (robotic arm) is asked to perform a high-level task with objects (eg stacking) present in a scene. The agent has to suggest a plan consisting of primitive actions (eg simple movement, grasping) in order to successfully complete the given task. Instructions are formulated in a way that forces the agent to perform visual reasoning over the presented scene before inferring the actions. We extend the recently introduced dataset SHOP-VRB with task instructions for each scene as well as an engine capable of assessing whether the sequence of primitives leads to a successful task completion. We also propose a novel approach based on multimodal attention for this task and demonstrate its performance on the new dataset.