课题基金 / 基金详情

Deep Spatiotemporal Models for Video Representation Learning

Deep Spatiotemporal Models for Video Representation Learning
用于视频表示学习的深度时空模型
批准号:
2431426
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
视频表示学习是人工智能社区感兴趣的领域,因为它旨在通过使用场景的动态来做出关键决策来促进计算机视觉领域的发展。与使用静态图像相比,视频数据为当前的信息提供了更多的背景,这提高了建立在视频表示学习框架上的系统做出决策的质量。该研究旨在通过开发一种基于图的机器学习模型来增加视频表示学习领域,该模型将优于当前最先进的模型。尽管视频数据不具有自然发生的图形结构(与社交网络不同),但使用基于图形的架构可以显著减少模型[1]中所需的参数数量。这使得所提出的模型适用于内存受限的设备,如移动电话Shirian, A., Tripathi, S., & Guha, T.(2021)。基于可学习图和图初始网络的动态情绪建模。IEEE多媒体汇刊。研究的目的和目标*研究的目的是:1。2.制作视频建模的时空图;学习邻接关系,使其与数据相关;说谎。扩展到异构图形,其中数据模式可以是多种(例如,视频和音频)。研究方法的新颖性(如果有的话)*该研究将有助于视频表示学习的现有知识体系。该研究旨在设计一种新的基于图的机器学习架构,新的损失函数来更好地惩罚我们的模型,以及优化技术来加速模型训练。潜在的影响、应用和益处*该研究将:改进现有的自主导航系统;2 .应用于监控系统中的目标检测和动作预测;3 .用于提高机器人的计算机视觉;适用于手机等内存受限的设备;能够叙述场景中发生的事件。这项研究与研究范围的关系*这项研究跨越了人工智能和机器人、数学科学以及信息和通信技术(ICT)等领域。这些都是EPSRC感兴趣的关键领域,因为这项研究将提高机器人的视觉感知能力,拓宽图论在社交网络之外的应用知识,而且自动驾驶汽车的未来也不远了。研究范畴;信息和通信技术,数学科学外部合作伙伴-英特尔实验室,圣地亚哥。
英文摘要
The area of video representation learning is of interest to the Artificial Intelligence community, as it aims to foster the field of computer vision by using the dynamics of a scene to make critical decisions. Compared to using still images, video data gives more context to the information present, and this improves the quality of decisions made by systems built on the framework of video representation learning. The research aims to add to the field of video representation learning by developing a graph-based machine learning model that would outperform current state-of-the-art models. Although video data do not have a naturally occurring graph structure (unlike social networks), using a graph-based architecture significantly reduces the number of parameters needed in the model [1]. And this makes the proposed model suitable for memory-constrained devices like mobile phones.[1] Shirian, A., Tripathi, S., & Guha, T. (2021). Dynamic Emotion Modeling with Learnable Graphs and Graph Inception Network. IEEE Transactions on Multimedia.The aims and objectives of the research *The objective of the research is to:1. develop spatiotemporal graphs for modeling videos;2. learn the adjacency such that it is data dependent; and3. extend to heterogeneous graphs where data modalities can be multiple (e.g.,video with audio).The novelty of the research methodology (if any) *The research would contribute to the existing body of knowledge in video representation learning. The research aims to design a new graph-based machine learning architecture, new loss functions to better penalize our model and optimization techniques to speed up model training.The potential impact, applications, and benefits *The research would:1. improve the current autonomous navigation systems;2. be applied for object detection and action prediction in surveillance systems;3. be used to improve computer vision in robots;4. be suitable for memory-constrained devices like mobile phones;5. be able to narrate events happening in a scene. And this is useful for visuallyimpaired individuals etc.How the research relates to the remit *The research cuts across the field of Artificial Intelligence and robotics, mathematical science, and Information and communication technologies (ICT). And these are key areas of interest for the EPSRC, as this research wouldimprove visual perception in robotics, broaden the knowledge on the applicability of graph theory beyond social networks and the future of self driving cars would not be far from reach. Research Category; ICT [Information and Communication Technologies], Mathematical SciencesExternal Partner - Intel Labs, San Diego.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于分子动力学的沥青/集料界面行为Spatiotemporal模型
  • 批准号:
    51378073
  • 项目类别:
    面上项目
  • 资助金额:
    72.0万元
  • 批准年份:
    2013
  • 负责人:
    裴建中
  • 依托单位: