课题基金 / 基金详情

CAREER: Visual Question Answering (VQA)

CAREER: Visual Question Answering (VQA)
职业:视觉问答 (VQA)
批准号:
1552377
负责人:
Devi Parikh
金额:
$51.7万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2017-01-31

项目摘要

项目成果

Devi Parikh的其他基金

相似基金

相关文献

中文摘要
翻译
该项目解决了视觉问答(VQA)的问题。给定一幅图像和关于该图像的自由形式的自然语言问题(例如,“这是哪种商店?”、“有多少人在排队?”、“过马路安全吗?”),机器的任务是自动产生一个简洁、准确、自由形式的自然语言答案(“面包房”、“5”、“是”)。VQA直接适用于各种具有高社会影响的应用程序,这些应用程序涉及人类从视觉数据中获取与情境相关的信息;在这些应用中,人和机器必须协作才能从图片中提取信息。例如,帮助视障用户了解他们的周围环境,分析人员根据大量监视做出决定,以及与机器人互动。该项目有可能从根本上改善视障用户的日常生活方式,并彻底改变整个社会与视觉数据的互动方式。该研究使得VQA代表的不是单一的狭义问题(如图像分类),而是丰富的语义场景理解问题和相关的研究方向。VQA中的每个问题都可能位于这个光谱的不同点上:从直接映射到现有的研究得很好的计算机视觉问题的问题(“这个房间叫什么?”=室内场景识别),一直到需要基于知识库的视觉(场景)、语言(语义)和推理(理解)的综合方法的问题(“后排可乐瓶旁边的披萨看起来像素食吗?”)。因此,这项工作映射到这一频谱上的一系列路点。在从不同角度解决VQA的动机下,这项研究计划正在产生新的数据集、知识和技术,包括(I)纯计算机视觉(Ii)整合视觉+语言(Iii)整合视觉+语言+常识(Iv)建立可解释的模型和(V)结合各种方法。此外,在以下方面正在做出新的贡献:(A)训练机器保持好奇心并主动提出问题进行学习(B)使用VQA作为一种方式来了解更多关于视觉世界的信息,而不是现有的注释方式所允许的;以及(C)训练机器知道它知道什么和不知道什么。
英文摘要
This project addresses the problem of Visual Question Answering (VQA). Given an image and a free-form natural language question about the image (e.g., "What kind of store is this?", "How many people are waiting in the queue?", "Is it safe to cross the street?"), the machine's task is to automatically produce a concise, accurate, free-form, natural language answer ("bakery", "5", "Yes"). VQA is directly applicable to a variety of applications of high societal impact that involve humans eliciting situationally-relevant information from visual data; where humans and machines must collaborate to extract information from pictures. Examples include aiding visually-impaired users in understanding their surroundings, analysts in making decisions based on large quantities of surveillance, and interacting with a robot. This project has the potential to fundamentally improve the way visually-impaired users live their daily lives, and revolutionize how society at large interacts with visual data. This research enables that VQA represents not a single narrowly-defined problem (e.g., image classification) but rather a rich spectrum of semantic scene understanding problems and associated research directions. Each question in VQA may lie at a different point on this spectrum: from questions that directly map to existing well-studied computer-vision problems ("What is this room called?" = indoor scene recognition) all the way to questions that require an integrated approach of vision (scene), language (semantics), and reasoning (understanding) over a knowledge base ("Does the pizza in the back row next to the bottle of Coke seem vegetarian?"). Consequently, this work maps to a sequence of waypoints along this spectrum. Motivated by addressing VQA from a variety of perspectives, this research program is generating new datasets, knowledge, and techniques in (i) pure computer vision (ii) integrating vision + language (iii) integrating vision + language + common sense (iv) building interpretable models and (v) combining a portfolio of methods. In addition, novel contributions are being made to (a) training the machine to be curious and actively ask questions to learn (b) using VQA as a modality to learn more about the visual world than what existing annotation modalities allow and (c) training the machine to know what it knows and what it does not.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Visual Question Answering (VQA)
  • 批准号:
    1661374
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $51.7万
  • 财政年份:
    2016
  • 负责人:
    Devi Parikh
  • 依托单位:
RI: Small: Debugging Machine Visual Recognition via Humans in the Loop
RI: Small: Debugging Machine Visual Recognition via Humans in the Loop
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
  • 批准号:
    60373031
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2003
  • 负责人:
    陈越
  • 依托单位: