BIGDATA: F: Audio-Visual Scene Understanding
BIGDATA: F: Audio-Visual Scene Understanding
批准号:
1741472
负责人:
Chenliang Xu
金额:
$65.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2022-08-31
中文摘要
理解我们周围的场景,即识别物体、人类行为和事件,并推断它们的空间、时间、关联和因果关系,是人类智能的一项基本能力。同样,设计能够理解场景的计算机算法是人工智能的一个基本问题。人类有意识或无意识地使用所有五种感官(视觉、听觉、味觉、嗅觉和触觉)来理解场景,因为不同的感官提供互补的信息。例如,看一部静音的电影会让人很难理解电影内容;在没有其他指引的情况下闭着眼睛走在街上是很危险的。然而,现有的机器场景理解算法仅依赖于单一模态。以视觉和听觉这两种最常用的感官为例,有针对每种单一模态的场景理解算法。然而,没有进行系统的调查来整合这两种模式,以实现更全面的视听场景理解。设计联合模拟音频和视觉模式的算法以实现完整的视听场景理解是很重要的,不仅因为这是人类理解场景的方式,而且因为它将在许多领域实现新的应用。这些领域包括多媒体(视频索引和场景编辑)、医疗保健(视觉和听觉受损人士的辅助设备)、监视安全(对可疑活动的全面监控)以及虚拟和增强现实(视觉和/或音轨的生成和替换)。此外,研究人员将让研究生和本科生参与研究活动,将研究成果纳入教学课程,并在当地学校和社区开展外展活动,目的是让更多人参与计算机科学。本项目旨在通过对互联网视频的大数据分析,克服单一模态方法的局限性,实现类人的视听场景理解。核心思想是学习将场景解析成元素,并推断出元素之间的关系,即形成一个视听场景图。具体地说,当事件显示相关的音频和视觉特征时,视听场景的元素可以是事件的联合视听组件。如果事件只以一种模式出现,它也可以是音频组件或视觉组件。要素之间的关系在较低层次上包括时空关系,在较高层次上包括相关关系和因果关系。通过这个场景图,可以提取、交换和解释两种模式之间的信息。研究者提出了三个主要的研究重点:(1)学习场景元素的联合视听表征;(2)学习场景图,组织场景元素;(3)跨模态场景补全。这三个研究重点都探索了视听场景理解空间的一个维度,但它们也是相互联系的。例如,视听场景元素是场景图中的节点,场景图又以结构化信息引导场景元素之间关系的学习;跨模态场景补全会在场景图中产生缺失数据,这对于良好的场景视听理解是必要的。本次提案的预期成果包括:一个学习各种场景元素的联合视听表示的软件包;一个网络部署的视听场景理解系统,利用学习到的场景元素和场景图,用文本生成来说明;基于场景理解的跨模态场景补全软件包以及一个大型视频数据集,带有用于视听关联、文本生成和场景补全的注释。数据集、软件和演示将放在项目网站上。
英文摘要
Understanding scenes around us, i.e., recognizing objects, human actions and events, and inferring their spatial, temporal, correlative and causal relations, is a fundamental capability in human intelligence. Similarly, designing computer algorithms that can understand scenes is a fundamental problem in artificial intelligence. Humans consciously or unconsciously use all five senses (vision, audition, taste, smell, and touch) to understand a scene, as different senses provide complimentary information. For example, watching a movie with the sound muted makes it very difficult to understand the movie; walking on a street with eyes closed without other guidance can be dangerous. Existing machine scene understanding algorithms, however, are designed to rely on just a single modality. Take the two most commonly used senses, vision and audition, as an example, there are scene understanding algorithms designed to deal with each single modality. However, no systematic investigations have been conducted to integrate these two modalities towards more comprehensive audio-visual scene understanding. Designing algorithms that jointly model audio and visual modalities towards a complete audio-visual scene understanding is important, not only because this is how humans understand scenes, but also because it will enable novel applications in many fields. These fields include multimedia (video indexing and scene editing), healthcare (assistive devices for visually and aurally impaired people), surveillance security (comprehensive monitoring of the suspicious activities), and virtual and augmented reality (generation and alternation of visuals and/or sound tracks). In addition, the investigators will involve graduate and undergraduate students in the research activities, integrate research results into the teaching curriculum, and conduct outreach activities to local schools and communities with an aim to broader participation in computer science. This project aims to achieve human-like audio-visual scene understanding that overcomes the limitations of single-modality approaches through big data analysis of Internet videos. The core idea is to learn to parse a scene into elements and infer their relations, i.e., forming an audio-visual scene graph. Specifically, an element of the audio-visual scene can be a joint audio-visual component of an event when the event shows correlated audio and visual features. It can also be an audio component or a visual component if the event only appears in one modality. The relations between the elements include spatial and temporal relations at a lower level, as well as correlative and causal relations at a higher level. Through this scene graph, information across the two modalities can be extracted, exchanged and interpreted. The investigators propose three main research thrusts: (1) Learning joint audio-visual representations of scene elements; (2) Learning a scene graph to organize scene elements; and (3) Cross-modality scene completion. Each of the three research thrusts explores a dimension in the space of audio-visual scene understanding, yet they are also inter-connected. For example, the audio-visual scene elements are nodes in the scene graph, and the scene graph, in turn, guides the learning of relations among scene elements with structured information; the cross-modality scene completion generates missing data in the scene graph and is necessary for good audio-visual understanding of the scene. Expected outcomes of this proposal include: a software package for learning joint audio-visual representations of various scene elements; a web-deployed system for audio-visual scene understanding utilizing the learned scene elements and scene graphs, illustrated with text generation; a software package for cross-modality scene completion based on scene understanding; and a large-scale video dataset with annotations for audio-visual association, text generation and scene completion. Datasets, software and demos will be hosted on the project website.
期刊论文(49)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2019
期刊:
IEEE Computer Society Conference on Computer Vision and Pattern Recognition workshops
影响因子:
--
作者:
[Tian, Yapeng, Shi, Jing, Li, Bochen, Duan, Zhiyao, Xu, Chenliang]
通讯作者:
Xu, Chenliang
DOI:
10.1109/lsp.2021.3076358
发表时间:
2020-10
期刊:
IEEE Signal Processing Letters
影响因子:
3.9
作者:
[You Zhang;Fei Jiang;Z. Duan]
通讯作者:
You Zhang;Fei Jiang;Z. Duan
DOI:
10.1109/tmm.2021.3099900
发表时间:
2021-07-26
期刊:
IEEE TRANSACTIONS ON MULTIMEDIA
影响因子:
7.3
作者:
[Eskimez, Sefik Emre, Zhang, You, Duan, Zhiyao]
通讯作者:
Duan, Zhiyao
DOI:
10.1109/wacv48630.2021.00117
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
作者:
[Shaojie Wang;Wentian Zhao;Ziyi Kou;Jing Shi;Chenliang Xu]
通讯作者:
Shaojie Wang;Wentian Zhao;Ziyi Kou;Jing Shi;Chenliang Xu
DOI:
10.1109/taslp.2019.2947741
发表时间:
2020-05
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[S. Eskimez;R. Maddox;Chenliang Xu;Z. Duan]
通讯作者:
S. Eskimez;R. Maddox;Chenliang Xu;Z. Duan
共 40 条
III: Small: Collaborative Research: Scalable Deep Bayesian Tensor Decomposition
-
批准号:1909912
-
项目类别:Standard Grant
-
资助金额:$19.91万
-
财政年份:2019
-
负责人:Chenliang Xu
-
依托单位:
RI: Small: Learning Dynamics and Evolution towards Cognitive Understanding of Videos
-
批准号:1813709
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2018
-
负责人:Chenliang Xu
-
依托单位:
海外基金