CAREER: Visual Question Answering (VQA)
CAREER: Visual Question Answering (VQA)
批准号:
1661374
负责人:
Devi Parikh
金额:
$51.7万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-10-01 至 2022-07-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This project addresses the problem of Visual Question Answering (VQA). Given an image and a free-form natural language question about the image (e.g., "What kind of store is this?", "How many people are waiting in the queue?", "Is it safe to cross the street?"), the machine's task is to automatically produce a concise, accurate, free-form, natural language answer ("bakery", "5", "Yes"). VQA is directly applicable to a variety of applications of high societal impact that involve humans eliciting situationally-relevant information from visual data; where humans and machines must collaborate to extract information from pictures. Examples include aiding visually-impaired users in understanding their surroundings, analysts in making decisions based on large quantities of surveillance, and interacting with a robot. This project has the potential to fundamentally improve the way visually-impaired users live their daily lives, and revolutionize how society at large interacts with visual data. This research enables that VQA represents not a single narrowly-defined problem (e.g., image classification) but rather a rich spectrum of semantic scene understanding problems and associated research directions. Each question in VQA may lie at a different point on this spectrum: from questions that directly map to existing well-studied computer-vision problems ("What is this room called?" = indoor scene recognition) all the way to questions that require an integrated approach of vision (scene), language (semantics), and reasoning (understanding) over a knowledge base ("Does the pizza in the back row next to the bottle of Coke seem vegetarian?"). Consequently, this work maps to a sequence of waypoints along this spectrum. Motivated by addressing VQA from a variety of perspectives, this research program is generating new datasets, knowledge, and techniques in (i) pure computer vision (ii) integrating vision + language (iii) integrating vision + language + common sense (iv) building interpretable models and (v) combining a portfolio of methods. In addition, novel contributions are being made to (a) training the machine to be curious and actively ask questions to learn (b) using VQA as a modality to learn more about the visual world than what existing annotation modalities allow and (c) training the machine to know what it knows and what it does not.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Visual Question Answering (VQA)
-
批准号:1552377
-
项目类别:Continuing Grant
-
资助金额:$51.7万
-
财政年份:2016
-
负责人:Devi Parikh
-
依托单位:
RI: Small: Debugging Machine Visual Recognition via Humans in the Loop
-
批准号:1341772
-
项目类别:Standard Grant
-
资助金额:$3.05万
-
财政年份:2013
-
负责人:Devi Parikh
-
依托单位:
RI: Small: Debugging Machine Visual Recognition via Humans in the Loop
-
批准号:1115719
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Devi Parikh
-
依托单位:
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
-
批准号:60373031
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2003
-
负责人:陈越
-
依托单位: