EAGER: Spatiotemporal Transformer for Activity Recognition
EAGER: Spatiotemporal Transformer for Activity Recognition
批准号:
2322993
负责人:
Scott Acton
金额:
$28.08万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2025-06-30
中文摘要
从视频中了解人类活动对于安全、国防、医学、机器人、制造业和教育等领域的多个应用非常重要。计算机视觉领域探索使用相机和计算机来自动执行任务,例如对象识别和活动识别。传统上,研究人员通过提取图像或视频中的组成特征并将这些特征与更复杂对象的模型进行匹配来开发计算机视觉系统。最近,已经应用了机器学习方法来训练计算机从数据而不是物理模型执行这样的识别任务。该项目探索了一种基于学习的对象识别方法,该方法基于学习视频中观察到的对象和人之间的语义关系。具体地说,该项目试图设计计算方法,该方法将自动导出数字视频中人与对象之间的关系,然后利用这些相关关系对人类行为进行分类(例如,踢球或握手)。与为理解语言而开发的机器学习方法不同,所提出的解决方案将使用特定于理解人类动作的视频元素,例如图像对象的检测,对象的运动以及视频中的空间和时间位置。计算机视觉解决方案的成功实施将允许自动分析视频中的人类活动。分析将有利于关键任务,如学习执行手术或理解在有效课堂中采取的行动。变压器是一种神经网络,它使用注意力来计算句子或一系列句子中单词之间的关系。Transformer模型的优点包括在长的单词序列上评估这些关系的能力,以及通过位置编码同时自动处理所有单词的能力。本项目不是采用为自然语言开发的Transformer并将其应用于视频问题,而是从第一原理开发视频Transformer。该系统的实现涉及机器学习设计中的三个明显进步。首先,所提出的方法通过用时间编码的光流信息将运动的概念作为特征引入到Transformer。其次,该方法允许几何和动作语义的运动特征之间的相互作用,利用在一个Transformer框架。最后,基于分布的注意力模型超越了传统的相关注意力概念。所提出的注意力模型捕捉动作序列中的显著相关性。这三个理论贡献加在一起,有可能大大推进视频理解。该奖项反映了NSF的法定使命,并已被认为值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估的支持。
英文摘要
Understanding human activity from video is important to several applications in security, defense, medicine, robotics, manufacturing, and education. The field of computer vision explores the use of cameras and computers to automate tasks such as object recognition and activity recognition. Traditionally, researchers have developed computer vision systems by extracting the constituent features in an image or video and matching those features to models of more complex objects. More recently, machine learning methods have been applied that train a computer to perform such a recognition task from data rather than a physical model. This project explores a learning-based object recognition approach based on learning semantic relationships between objects and people observed in video. Specifically, the project attempts to design computing methods that will automatically derive relationships between people and objects in digital video and then exploit those correlative relationships in classifying a human action (e.g., kicking a ball or shaking hands). Unlike machine learning methods developed for understanding language, the proposed solution will use elements of the video specific to understanding human action such as detection of imaged objects, the motion of objects, and the spatial and temporal position in the video. Successful implementation of the computer vision solution will allow human activities in video to be automatically analyzed. The analysis will benefit critical tasks such as learning to perform a surgery or understanding the actions taken in an effective classroom.Transformers are a type of neural network that use attention to compute relationships between words in a sentence or series of sentences. The advantages of the transformer model include the ability to assess these relationships over long sequences of words and the ability to automatically process all words simultaneously via a positional encoding. Instead of taking the transformer developed for natural language and fitting it to a video problem, this project seeks to develop a video transformer from first principles. The realization of this system involves three distinct advances in the machine learning design. First, the proposed approach brings the concept of motion as a feature to the transformer by way of optical flow information encoded with time. Second, the proposed method allows interactions between geometric and motion features of action semantics to exploited in a transformer framework. Last, the distribution-based attention model goes beyond the traditional correlative notion of attention. The proposed attention model captures significant correlations in action sequences. Together, the three theoretical contributions have the potential to significantly advance video understanding.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Intergovernmental Personnel Act Assignment
-
批准号:1950730
-
项目类别:Intergovernmental Personnel Award
-
资助金额:$26.47万
-
财政年份:2019
-
负责人:Scott Acton
-
依托单位:
ABI Innovation: Towards the Neurome -- Automated Image Analysis for Neuroinformatics
-
批准号:1062433
-
项目类别:Standard Grant
-
资助金额:$48.35万
-
财政年份:2011
-
负责人:Scott Acton
-
依托单位:
Decentralized Image Retrieval for Education (DIRECT)
-
批准号:0121596
-
项目类别:Standard Grant
-
资助金额:$49.44万
-
财政年份:2002
-
负责人:Scott Acton
-
依托单位:
国内基金
海外基金
基于分子动力学的沥青/集料界面行为Spatiotemporal模型
-
批准号:51378073
-
项目类别:面上项目
-
资助金额:72.0万元
-
批准年份:2013
-
负责人:裴建中
-
依托单位: