III: Medium: Learning Multimodal Knowledge about Entities and Events
III: Medium: Learning Multimodal Knowledge about Entities and Events
批准号:
1703166
负责人:
Hanna Hajishirzi
金额:
$70.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2022-07-31
中文摘要
关于世界的日常知识是智能信息处理和推理的必要条件。人们可以读懂文字的字里行间,看到图像中看不到的东西,因为人们每天都有关于世界如何运行的功能性知识。本研究的主要目标是开发学习算法,能够从大规模的多模式网络数据中自动获取以实体和事件为中心的知识。实体知识包括关于对象和人的广泛的物理和概念知识,包括它们的属性、它们的相对差异以及它们之间的逻辑关系。事件知识是通过子事件和事件参与者之间的层次和时间关系组织起来的关于人们生活中日常事件的结构化知识。总而言之,由此产生的知识将是向前迈出的关键一步,使强大的人工智能系统能够在自然语言处理和计算机视觉的交叉点上能够理解非结构化多模式信息并进行推理。这项研究的潜在影响包括为视障人士提供的交互式辅助系统和多模式教育界面。本项目研究多通道知识提取作为一种新的研究范式,将自然语言处理中的相关方法如信息提取、文本蕴涵和框架语义与计算机视觉的最新进展联系起来。常识知识获取的关键挑战之一是克服报道偏差,即人们没有陈述显而易见的事情。因此,该项目开发了基于基于图形的集体推理的新学习算法,该集体推理可以对系统地影响人们用语言、图像和视频描述世界的方式的未说出的知识进行推理。此外,该项目还开发了用于视觉语义解析和事件识别的新模型,该模型通过指定事件的各种结构组件(如参与者、对象、位置、工具、意图和目标)来概括现有的活动识别研究。学习到的知识和表达将通过几个应用程序进行验证,包括多模式问题回答和扎根语言理解。
英文摘要
Everyday knowledge about the world is a necessary condition for intelligent information processing and reasoning. People can read between the lines in text and see beyond what are visible in images because of everyday functional knowledge about how the world works. The primary goal of this research is to develop learning algorithms that can automatically acquire such knowledge, centered around entities and events, from large-scale multimodal web data. Entity knowledge includes a broad range of physical and conceptual knowledge about objects and people, including their attributes, their relative differences, and logical relations among them. Event knowledge focuses on structural knowledge about everyday events in people's lives organized through hierarchical and temporal relations among sub-events and the event participants. Together, the resulting knowledge will be a critical step forward to enable robust AI systems at the intersection between natural language processing and computer vision that can understand and reason about unstructured multimodal information. The potential impact of this research includes interactive assistive systems for the visually-impaired and multimodal educational interfaces. This project investigates multimodal knowledge extraction as a new research paradigm drawing connections between relevant methods in natural language processing such as information extraction, textual entailments, and frame semantics with recent advances in computer vision. One of the critical challenges in commonsense knowledge acquisition is to overcome reporting bias, i.e., people do not state the obvious. Therefore, this project develops new learning algorithms based on a graph-based collective inference that can reason about unspoken knowledge that systematically influences the way people describe the world in language, images, and videos. In addition, this project develops new models for visual semantic parsing and event recognition, which generalize existing studies on activity recognition by specifying various structural components of events such as actors, objects, locations, tools, intents, and goals. The learned knowledge and representation will be validated through several applications including multimodal question answering and grounded language understanding.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Knowledge-Rich Neural Text Comprehension and Reasoning
-
批准号:2044660
-
项目类别:Continuing Grant
-
资助金额:$54.98万
-
财政年份:2021
-
负责人:Hanna Hajishirzi
-
依托单位:
IIS: RI: Travel Proposal: Student Travel Support for the 2019 Association for Computational Linguistics Student Research Workshop
-
批准号:1929269
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Hanna Hajishirzi
-
依托单位:
RI: Small: Learning to Read, Ground, and Reason in Multimodal Text
-
批准号:1616112
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2016
-
负责人:Hanna Hajishirzi
-
依托单位:
EAGER: Generating and Understanding Narratives for Dynamic Environments
-
批准号:1352249
-
项目类别:Standard Grant
-
资助金额:$14.99万
-
财政年份:2013
-
负责人:Hanna Hajishirzi
-
依托单位:
海外基金