Computer Vision - ACCV 2020 - 15th Asian Conference on Computer Vision, Kyoto, Japan, November 30 - December 4, 2020, Revised Selected Papers, Part III
Computer Vision - ACCV 2020 - 15th Asian Conference on Computer Vision, Kyoto, Japan, November 30 - December 4, 2020, Revised Selected Papers, Part III
复制标题
计算机视觉 - ACCV 2020 - 第 15 届亚洲计算机视觉会议,日本京都,2020 年 11 月 30 日至 12 月 4 日,修订后的精选论文,第三部分
DOI:
10.1007/978-3-030-69535-4_44
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Nazarczuk M
中科院分区:
文献类型:
--
作者:
Nazarczuk M
In this work, we present a new AI task-Vision to Action (V2A)-where an agent (robotic arm) is asked to perform a high-level task with objects (eg stacking) present in a scene. The agent has to suggest a plan consisting of primitive actions (eg simple movement, grasping) in order to successfully complete the given task. Instructions are formulated in a way that forces the agent to perform visual reasoning over the presented scene before inferring the actions. We extend the recently introduced dataset SHOP-VRB with task instructions for each scene as well as an engine capable of assessing whether the sequence of primitives leads to a successful task completion. We also propose a novel approach based on multimodal attention for this task and demonstrate its performance on the new dataset.