Collaborative Research: RI:Medium:Understanding Events from Streaming Video - Joint Deep and Graph Representations, Commonsense Priors, and Predictive Learning
Collaborative Research: RI:Medium:Understanding Events from Streaming Video - Joint Deep and Graph Representations, Commonsense Priors, and Predictive Learning
批准号:
1955230
负责人:
Sathyanarayanan Aakur
金额:
$28.51万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2023-11-30
中文摘要
虽然人类很容易处理视频数据并从中提取意义,但设计这样做的算法却非常困难。一旦开发出来,这项技术有很多应用,比如建造辅助机器人或建造独立生活的智能空间或监测野生动物。视频数据捕获事件,这是人类经验内容的核心。事件由物体/人(谁)、地点(何处)、时间(何时)、行为(什么)、活动(如何)和意图(为什么)组成。该项目开发了一种基于计算机视觉的事件理解算法,该算法以自监督的流方式运行。该算法将预测和检测新旧事件,学习构建分层事件表示,所有这些都是在一个随时间更新的先验知识库的背景下进行的。其目的是产生对事件的解释,超越所见,而不仅仅是识别。本研究通过将自监督学习过程与先验知识相结合,将该领域推向开放世界算法,并且很少或不需要监督,从而推动了计算机视觉的前沿。此外,该项目将重点关注大一和大二本科女生的招募和保留,并关注三个地点的少数族裔学生:南佛罗里达大学、佛罗里达州立大学和俄克拉荷马州立大学。该方法的核心是混合表示层次结构,包括连续表示和基于符号图的表示。连续值表示是标准的、向量值深度学习堆栈,以知识库中某个对象或动作概念的嵌入向量结束。表征的下一个层次由这些动词和名词的基本符号组成。当这些基本组合与知识库中的概念相关联时,它们构成事件解释,包含超出图像中观察到的描述。这些符号层次是使用格伦纳德模式理论中的规范表示来构建的。这些表示具有灵活的图形结构骨架,比其他已知的图形模型更具表现力。该项目的具体技术目标有四个方面。首先,将模式理论中基于函数的连续和基于能量的Grenander正则符号表示整合为一个基于均衡传播的集成公式。其次,它将研究和开发使用和修改常识性知识库的方法。这将有助于超越封闭世界的假设,这在当前基于注释数据的深度学习方法的实践中是隐含的。第三,它将在图流形上开发动态模型,这将使图结构的生成建模能够用于预测和发现新概念。第四,受人类感知实验和神经科学发现的启发,它将在连续和符号表征上设计预测性自监督学习。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
While it is easy for humans to process video data and extract meanings from it, it is extremely hard to design algorithms to do so. When developed, there are many applications of this technology, such as building assistive robotics or constructing smart spaces for independent living or monitoring wildlife. Video-data capture events, which are central to the content of human experience. Events consist of objects/people (who), location (where), time (when), actions (what), activities (how), and intent (why). This project develops a computer vision-based event understanding algorithm that operates in a self-supervised, streaming fashion. The algorithm will predict and detect old and new events, learn to build hierarchical event representations, all in the context of a prior knowledge-base that is updated over time. The intent is to generate interpretations of an event that go beyond what is seen, rather than just recognition. This research pushes the frontier of computer vision by coupling the self-supervised learning process with prior knowledge, moving the field towards open-world algorithms, and needing little or no supervision. Furthermore, this project will focus on recruitment and retention of undergraduate women students through freshman and sophomore years, with attention towards underrepresented minority students at the three sites: University of South Florida, Florida State University, and Oklahoma State University.At the core of the approach is a hybrid representational hierarchy that includes both continuous representations and symbolic graph-based representations. The continuous-valued representation is the standard, vector-valued deep learning stack that ends in an embedding vector of some object or action concept in the knowledge base. The next level of the representation consists of elementary symbolic compositions of these verbs and nouns. These elementary compositions, when associated with concepts from a knowledge-base they makeup an event interpretation, containing descriptions that go beyond what is observed in the image. These symbolic levels are built using Grenander's canonical representations from pattern theory. These representations, which have flexible graph-structured backbones, are more expressive than other well-known graphical models. The specific technical aims of the project are four-fold. First, it seeks to integrate function-based continuous with energy-based Grenander's canonical symbolic representations from pattern theory into one integrated formulation based on equilibrium propagation. Second, it will research and develop ways to use and modify commonsense knowledge bases. This will help to go beyond the closed world assumption, which is implicit in the current practice of annotated data-based deep learning approaches. Third, it will develop dynamical models on graph manifolds, which will enable generative modeling of graph structures for prediction and discovery of new concepts. Fourth, inspired by finding from human perception experiments and neuroscience, it will design predictive self-supervised learning over both continuous and symbolic representations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/978-3-031-19833-5_26
发表时间:
2022
期刊:
影响因子:
--
作者:
[A. Bal;R. Mounir;Sathyanarayanan N. Aakur;Sudeep Sarkar;Anuj Srivastava]
通讯作者:
A. Bal;R. Mounir;Sathyanarayanan N. Aakur;Sudeep Sarkar;Anuj Srivastava
Action Localization Through Continual Predictive Learning
通过持续预测学习进行动作本地化
DOI:
10.1007/978-3-030-58568-6_18
发表时间:
2020
期刊:
European Conference on Computer Vision
影响因子:
--
作者:
[Aakur, Sathyanarayanan, Sarkar, Sudeep]
通讯作者:
Sarkar, Sudeep
Actor-centered Representations for Action Localization in Streaming Videos. European Conference on Computer Vision
流视频中动作本地化的以演员为中心的表示。
DOI:
--
发表时间:
2022
期刊:
European Conference on Computer Vision
影响因子:
--
作者:
[Aakur, Sathyanarayanan N., Sarkar, Sudeep]
通讯作者:
Sarkar, Sudeep
Unsupervised Gaze Prediction in Egocentric Videos by Energy-based Surprise Modeling
通过基于能量的惊喜建模在自我中心视频中进行无监督注视预测
DOI:
--
发表时间:
2021
期刊:
Imaging and Computer Graphics Theory and Applications
影响因子:
--
作者:
[Aakur, Sathyanarayanan N., Bagavathi, Arunkumar]
通讯作者:
Bagavathi, Arunkumar
DOI:
10.1109/icpr56361.2022.9956441
发表时间:
2022-08
期刊:
2022 26th International Conference on Pattern Recognition (ICPR)
影响因子:
--
作者:
[Priyadharsini Ramamurthy;Sathyanarayanan N. Aakur]
通讯作者:
Priyadharsini Ramamurthy;Sathyanarayanan N. Aakur
共 7 条
CAREER:Towards Causal Multi-Modal Understanding with Event Partonomy and Active Perception
-
批准号:2348690
-
项目类别:Continuing Grant
-
资助金额:$51.42万
-
财政年份:2023
-
负责人:Sathyanarayanan Aakur
-
依托单位:
Collaborative Research: RI:Medium:Understanding Events from Streaming Video - Joint Deep and Graph Representations, Commonsense Priors, and Predictive Learning
-
批准号:2348689
-
项目类别:Continuing Grant
-
资助金额:$28.51万
-
财政年份:2023
-
负责人:Sathyanarayanan Aakur
-
依托单位:
CAREER:Towards Causal Multi-Modal Understanding with Event Partonomy and Active Perception
-
批准号:2143150
-
项目类别:Continuing Grant
-
资助金额:$51.42万
-
财政年份:2022
-
负责人:Sathyanarayanan Aakur
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: