课题基金 / 基金详情

Developing Foundation Model Capabilities for Video Understanding in the Open World

Developing Foundation Model Capabilities for Video Understanding in the Open World
开发开放世界中视频理解的基础模型能力
批准号:
2711268
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目的目标是开发用于视频理解的开放世界深度学习模型,允许用户使用自然语言描述询问有关视频内容的查询。开发的方法将利用预先训练的基础模型,而不是从头开始训练深度学习模型。基础模型是在大量数据上训练的机器学习模型,可用于解决各种下游任务。扩展基础模型以解决特定问题通常需要比从头开始训练更少的数据,并提高了专家模型的通用性。将开发方法来解决开放世界的图像级问题,需要使用预先训练的基础模型进行自然语言输入。从这些发展中获得的见解将激发解决视频中类似问题的模型的构建。问题的示例包括计算图像中的文本指定对象和视频中的重复,以及回答有关场景中对象的区域,形状和结构的查询。而不是解决一个特定的类的问题,开发的模型将允许用户通过在推理时提供有关感兴趣的类的文本输入来解决任何任意类的问题。重要的是,使这种开放世界模型适应新的类不需要额外的训练或数据,即使在训练过程中类是不可见的。因此,这项工作将导致AI系统更容易被公众访问,他们可能无法访问大量的标记数据和计算通常需要训练特定于类的模型。开发图像理解模型,允许用户使用自然语言对图像内容提出问题。2.利用步骤(1)中的见解和方法,开发具有类似视频理解功能的模型,允许用户使用自然语言对视频内容提出问题。例如,在步骤(1)中开发用于使用文本对图像中的对象进行计数的模型可以启发模型在步骤(2)中使用文本对视频中的对象进行计数。重复步骤(1)和(2),添加更多功能。研究方法的新奇虽然利用预训练的视觉语言基础模型进行图像检索,对象检测和实例分割等任务已被显着探索图像,但类似的开发较少用于视频。这是因为由于额外的时间维度,从视频中学习更加复杂。此外,该项目中开发的方法将包括新的深度学习架构,这些架构更通用,在现有任务中表现更好,或者解决新问题,例如使用自然语言描述的视频中的重复计数,以及回答关于大小,形状,与EPSRC的战略和研究领域保持一致该项目涉及“人工智能技术”研究领域。是否有公司或合作者参与?号
英文摘要
DescriptionThe goal of this project is to develop open-world deep learning models for video understanding that allow users to ask queries about video content using natural language descriptions. Rather than training deep learning models from scratch, the methods developed will leverage pre-trained foundation models. A foundation model is a machine learning model trained on large quantities of data that can be adapted to solve a wide variety of downstream tasks. Extending a foundation model to solve a specific problem usually requires less data than training it from scratch and improves the generalizability of the specialist model. Methods will be developed to solve open-world image-level problems requiring natural language input with pre-trained foundation models. Insights from these developments will inspire the construction of models that solve analogous problems in videos. Examples of problems include counting text-specified objects in images and repetitions in videos and answering queries about the area, shape, and structure of objects in a scene. Rather than solving a problem for a particular class, the models developed will allow users to solve the problem for any arbitrary class by providing text input about the class of interest at inference time. Importantly, adapting such open-world models to new classes would require no additional training or data, even if the class were unseen during training. Hence, this work will result in AI systems that are more accessible to the general public, who may not have access to the large quantities of labelled data and compute typically necessary to train class-specific models.Aims & Objectives1. Develop models for image understanding that allow users to ask questions on the image content using natural language.2. Leverage insights and methods from step (1) to develop models with similar capabilities for video understanding that allow users to ask questions on video content using natural language. For instance, a model developed to count objects in images using text in step (1) could inspire a model to count objects in videos using text in step (2).3. Iterate on steps (1) and (2), adding more capabilities. Novelty of the Research MethodologyWhile leveraging pre-trained vision-language foundation models for tasks such as image retrieval, object detection, and instance segmentation has been significantly explored for images, similar developments have been less explored for videos. This is because learning from videos is more complex due to an additional temporal dimension. Furthermore, methods developed in this project will include novel deep learning architectures that are more general and perform better at existing tasks or that solve new problems such as repetition counting in videos using natural language descriptions and answering arbitrary natural language queries about the size, shape, and structure of objects.Alignment to the EPSRC's Strategies & Research AreasThis project relates to the "Artificial Intelligence Technologies" research area.Any Companies or Collaborators Involved?No.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金