课题基金 / 基金详情

CAREER: Structure-Preserving Multimodal Alignment between Vision and Language

CAREER: Structure-Preserving Multimodal Alignment between Vision and Language
职业:视觉和语言之间保持结构的多模态对齐
批准号:
2239840
负责人:
Humphrey Shi
金额:
$56.3万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2028-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
人工智能(AI)面临的一个巨大挑战是能够处理多模态视觉和语言数据,同时保持这些模态之间的关系,以便维持不同模态之间的联系。目前的机器学习系统并不能完全掌握人类视觉和语言中存在的结构和关系,因此在可解释性、效率、可测量性和因果关系方面难以产生预期的结果。该项目解决了机器学习中的基本多模态对齐问题,并将推进计算机视觉和自然语言处理的研究,特别是在多模态视觉语言生成和理解的颠覆性创新领域。它将导致视觉和语言的理论理解以及实际应用的突破。在这个项目下开发的技术可以类似地用于连接不同类型的潜在结构跨模态,并不限于视觉和语言。这对于科学领域负责任的人工智能应用非常有益,因为人们不仅想了解数据中的关系,还想了解结构和因果解释。这种理解对于减少机器学习模型表现出的人口统计学偏见也至关重要。通过教育、开源和外展活动,该项目将培训和教育从K-12到研究生的所有级别的学生,推进理论视野和语言课程,减少偏见,并进一步民主化AI。保护结构是理解如何使机器学习模型更好,更可靠的重要组成部分。该项目旨在通过保持结构的潜在空间对齐,在多模态视觉和语言建模方面创造新颖而重大的科学进步,以在视觉和语言之间建立桥梁。该项目旨在增加语言和视觉嵌入的结构保留性质,并在两种潜在表示之间建立一个映射,以保留底层结构。具体而言,该项目将通过四个方面来实现这些目标:(I)开发结构保持的潜在表征和视觉与语言之间的映射;(II)通过潜在结构提高学习和数据效率;(III)通过结构信息开发新的评估指标,以提高可测量性;(四)建立一个因果表述和解释框架。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值进行评估,更广泛的影响审查标准。
英文摘要
A grand challenge in artificial intelligence (AI) is to be able to process multimodal vision and language data, while preserving relationships across such modalities so that the linkages between the different modalities is sustained. Current machine learning systems do not fully grasp the structures and relationships that exist within human vision and language, and thus have difficulties producing the desired outcomes in terms of interpretability, efficiency, measurability, and causality. This project tackles the fundamental multimodal alignment problem in machine learning and will advance research in both computer vision and natural language processing, especially in the disruptive innovation areas of multimodal vision-language generation and understanding. It will lead to breakthroughs in both theoretical understanding as well as practical applications of vision and language. The techniques developed under this project could similarly be used to connect different types of latent structures across modalities and are not limited to vision and language. This would be extremely beneficial for responsible AI applications in the sciences, where people not only want to understand the relationship in data, but the structure and causal explanations. Such an understanding is also critical for reducing demographic biases that machine learning models exhibit. Through education, open-sourcing and outreach activities, this project will train and educate students of all levels - from K-12 to graduate - in AI, advance theoretical vision and language courses, reduce bias, and further democratize AI.Preserving structure is an essential component of understanding how to make machine learning models better and more reliable. This project aims to create novel and significant scientific advances in multimodal vision and language modeling with structure-preserving latent space alignment to build a bridge between vision and language. The project aims to increase the structural preserving nature for linguistic and visual embeddings and develop a map between the two latent representations that preserves the underlying structures. In particular, the project will achieve these goals through four thrusts: (I) Developing structure-preserving latent representations and mapping between vision and language; (II) Improving learning and data efficiency through latent structures; (III) Develop novel evaluation metrics through structural information to improve measurability; (IV) Develop a causal representation and interpretation framework.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/iccv51070.2023.00713
发表时间: 2022-11
期刊: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子: --
作者: [Xingqian Xu;Zhangyang Wang;Eric Zhang;Kai Wang;Humphrey Shi]
通讯作者: Xingqian Xu;Zhangyang Wang;Eric Zhang;Kai Wang;Humphrey Shi
DOI: 10.1109/iccv51070.2023.01462
发表时间: 2023-03
期刊: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子: --
作者: [Levon Khachatryan;A. Movsisyan;Vahram Tadevosyan;Roberto Henschel;Zhangyang Wang;Shant Navasardyan;Humphrey Shi]
通讯作者: Levon Khachatryan;A. Movsisyan;Vahram Tadevosyan;Roberto Henschel;Zhangyang Wang;Shant Navasardyan;Humphrey Shi
海外基金