课题基金 / 基金详情

EAGER: Spatiotemporal Transformer for Activity Recognition

EAGER: Spatiotemporal Transformer for Activity Recognition
EAGER:用于活动识别的时空转换器
批准号:
2322993
负责人:
Scott Acton
金额:
$28.08万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2025-06-30

项目摘要

项目成果

Scott Acton的其他基金

相似基金

相关文献

中文摘要
翻译
从视频中了解人类活动对于安全、国防、医疗、机器人、制造和教育中的几个应用程序非常重要。计算机视觉领域探索使用摄像机和计算机来自动执行诸如物体识别和活动识别之类的任务。传统上,研究人员通过提取图像或视频中的组成特征,并将这些特征与更复杂对象的模型进行匹配,来开发计算机视觉系统。最近,机器学习方法已经被应用,这些方法训练计算机从数据而不是物理模型执行这样的识别任务。这个项目探索了一种基于学习的对象识别方法,该方法基于学习视频中观察到的对象和人之间的语义关系。具体地说,该项目试图设计计算方法,自动得出数字视频中人和对象之间的关系,然后在对人类行为(例如踢球或握手)进行分类时利用这些相关关系。与为理解语言而开发的机器学习方法不同,所提出的解决方案将使用视频中特定于理解人类行为的元素,如检测成像的对象、对象的运动以及视频中的空间和时间位置。计算机视觉解决方案的成功实施将使视频中的人类活动能够被自动分析。这种分析将有助于关键任务,如学习进行手术或理解在有效课堂上采取的行动。变形金刚是一种神经网络,它使用注意力来计算句子或一系列句子中单词之间的关系。转换器模型的优点包括能够评估长序列单词上的这些关系,以及通过位置编码自动同时处理所有单词的能力。这个项目不是采用为自然语言开发的转换器并将其应用于视频问题,而是从基本原理出发开发视频转换器。该系统的实现涉及到机器学习设计的三个不同方面的进步。首先,该方法通过用时间编码的光流信息将运动的概念作为特征引入到变压器中。其次,该方法允许在变换框架中利用动作语义的几何特征和运动特征之间的交互。最后,基于分布的注意力模型超越了传统的相关注意力概念。所提出的注意力模型捕获了动作序列中的显著相关性。这三个理论贡献加在一起,有可能显著提高对视频的理解。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding human activity from video is important to several applications in security, defense, medicine, robotics, manufacturing, and education. The field of computer vision explores the use of cameras and computers to automate tasks such as object recognition and activity recognition. Traditionally, researchers have developed computer vision systems by extracting the constituent features in an image or video and matching those features to models of more complex objects. More recently, machine learning methods have been applied that train a computer to perform such a recognition task from data rather than a physical model. This project explores a learning-based object recognition approach based on learning semantic relationships between objects and people observed in video. Specifically, the project attempts to design computing methods that will automatically derive relationships between people and objects in digital video and then exploit those correlative relationships in classifying a human action (e.g., kicking a ball or shaking hands). Unlike machine learning methods developed for understanding language, the proposed solution will use elements of the video specific to understanding human action such as detection of imaged objects, the motion of objects, and the spatial and temporal position in the video. Successful implementation of the computer vision solution will allow human activities in video to be automatically analyzed. The analysis will benefit critical tasks such as learning to perform a surgery or understanding the actions taken in an effective classroom.Transformers are a type of neural network that use attention to compute relationships between words in a sentence or series of sentences. The advantages of the transformer model include the ability to assess these relationships over long sequences of words and the ability to automatically process all words simultaneously via a positional encoding. Instead of taking the transformer developed for natural language and fitting it to a video problem, this project seeks to develop a video transformer from first principles. The realization of this system involves three distinct advances in the machine learning design. First, the proposed approach brings the concept of motion as a feature to the transformer by way of optical flow information encoded with time. Second, the proposed method allows interactions between geometric and motion features of action semantics to exploited in a transformer framework. Last, the distribution-based attention model goes beyond the traditional correlative notion of attention. The proposed attention model captures significant correlations in action sequences. Together, the three theoretical contributions have the potential to significantly advance video understanding.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Intergovernmental Personnel Act Assignment
  • 批准号:
    1950730
  • 项目类别:
    Intergovernmental Personnel Award
  • 资助金额:
    $26.47万
  • 财政年份:
    2019
  • 负责人:
    Scott Acton
  • 依托单位:
ABI Innovation: Towards the Neurome -- Automated Image Analysis for Neuroinformatics
  • 批准号:
    1062433
  • 项目类别:
    Standard Grant
  • 资助金额:
    $48.35万
  • 财政年份:
    2011
  • 负责人:
    Scott Acton
  • 依托单位:
Decentralized Image Retrieval for Education (DIRECT)
  • 批准号:
    0121596
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.44万
  • 财政年份:
    2002
  • 负责人:
    Scott Acton
  • 依托单位:
国内基金
海外基金
基于分子动力学的沥青/集料界面行为Spatiotemporal模型
  • 批准号:
    51378073
  • 项目类别:
    面上项目
  • 资助金额:
    72.0万元
  • 批准年份:
    2013
  • 负责人:
    裴建中
  • 依托单位: