Embodied Question Answering

Embodied Question Answering
复制标题

DOI:
10.1109/cvprw.2018.00279
复制
发表时间:
2018-06
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Abhishek Das;Samyak Datta;Georgia Gkioxari;Stefan Lee;Devi Parikh;Dhruv Batra
Abhishek Das;Samyak Datta;Georgia Gkioxari;Stefan Lee;Devi Parikh;Dhruv Batra
中科院分区:
其他
文献类型:
--
作者:
Abhishek Das;Samyak Datta;Georgia Gkioxari;Stefan Lee;Devi Parikh;Dhruv Batra

文献摘要

被引文献

相似文献

我们提出了一个新的人工智能任务--智能问答(Question Questioning,简称QA)--在3D环境中的随机位置产生一个智能体,并提出一个问题(“汽车是什么颜色的?').为了回答,智能体必须首先智能地导航以探索环境,通过第一人称(自我中心)视觉收集必要的视觉信息,然后回答问题(“橙色”)。QQA需要一系列AI技能-语言理解,视觉识别,主动感知,目标驱动导航,常识推理,长期记忆,以及将语言转化为行动。在这项工作中,我们开发了一个House 3D环境中的问题和答案数据集[1],评估指标,以及一个经过模仿和强化学习训练的分层模型。
We present a new AI task - Embodied Question Answering(EmbodiedQA) - where an agent is spawned at a random location in a 3D environment and asked a question ('What color is the car?'). In order to answer, the agent must first intelligently navigate to explore the environment, gather necessary visual information through first-person (egocentric) vision, and then answer the question ('orange'). EmbodiedQA requires a range of AI skills - language understanding, visual recognition, active perception, goal-driven navigation, commonsense reasoning, long-term memory, and grounding language into actions. In this work, we develop a dataset of questions and answers in House3D environments [1], evaluation metrics, and a hierarchical model trained with imitation and reinforcement learning.