RI: Medium: Improving grounding, generalization and contextual reasoning in vision and language models
RI: Medium: Improving grounding, generalization and contextual reasoning in vision and language models
批准号:
2107048
负责人:
Olga Russakovsky
金额:
$120.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-01 至 2025-08-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Recent Artificial Intelligence (AI) advances have brought us closer to the possibility of important and exciting real-world applications: ranging from robot assistants for the elderly or differently-abled, to large-scale video analysis of footage from police body-worn cameras to examine police-civilian interactions. Such applications require AI models to understand both visual and natural language cues. However, the state of vision-and-language technology is still not quite ready for these scenarios. Current visual recognition models appear to recognize many different objects but lack an understanding of the interconnection and structure of the visual world. Current image captioning systems output reasonable but completely generic image descriptions. Modern visual question answering systems are not robust to simple changes like synonyms or word rearrangements. This research will lead to fundamental advances in visual recognition and natural language understanding, laying the groundwork for more effective human-machine collaboration. The goal of this research is to move towards a tighter, more accurate and contextual integration of visual recognition and natural language processing. This involves addressing three key challenges: (1) enabling accurate and scalable grounding by establishing robust bi-directional connections between visual input and natural language tokens; (2) improving generalization of vision-and-language models to novel concepts and tasks; and (3) enabling contextual reasoning to allow models to effectively adapt to human or task-specific needs. The unifying theme is that all three challenges require innovation in not only modeling but also in reliable and insightful benchmarking: current evaluation frameworks are insufficient to drive progress in this space. The roadmap is to redesign existing benchmarks and evaluation paradigms, use the newly formulated metrics to identify the shortcomings in existing systems, and rely on these insights to drive the deep learning modeling innovations. This research uses the team’s expertise in designing multi-modal models for vision and language as well as in constructing effective large-scale benchmarks. The findings will be disseminated through technical workshops, open access publications, and open-source code. They will also be integrated into undergraduate, graduate and K-12 curriculum through collaboration with foundations like AI4ALL.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2206.02916
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
作者:
[Zhiwei Deng-;Olga Russakovsky]
通讯作者:
Zhiwei Deng-;Olga Russakovsky
DOI:
10.48550/arxiv.2207.01206
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
作者:
[Shunyu Yao;Howard Chen;John Yang;Karthik Narasimhan]
通讯作者:
Shunyu Yao;Howard Chen;John Yang;Karthik Narasimhan
DOI:
10.1007/978-3-031-19781-9_14
发表时间:
2022-01
期刊:
影响因子:
--
作者:
[Zeyu Wang;Yu Wu;Karthik Narasimhan;Olga Russakovsky]
通讯作者:
Zeyu Wang;Yu Wu;Karthik Narasimhan;Olga Russakovsky
ReAct: Synergizing Reasoning and Acting in Language Models
ReAct:在语言模型中协同推理和行动
DOI:
--
发表时间:
2023
期刊:
International Conference on Learning Representations (ICLR
影响因子:
--
作者:
[Yao, Shunyu, Zhao, Jeffrey, Yu, Dian, Du, Nan, Shafran, Izhak, Narasimhan, Karthik, Cao, Yuan]
通讯作者:
Cao, Yuan
CAREER: Overcoming bias in computer vision: Building fairer systems and training diverse leaders
-
批准号:2145198
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2022
-
负责人:Olga Russakovsky
-
依托单位:
海外基金