课题基金 / 基金详情

Developing Foundation Model Capabilities for Video Understanding in the Open World

Developing Foundation Model Capabilities for Video Understanding in the Open World
开发开放世界中视频理解的基础模型能力
批准号:
2711268
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
DescriptionThe goal of this project is to develop open-world deep learning models for video understanding that allow users to ask queries about video content using natural language descriptions. Rather than training deep learning models from scratch, the methods developed will leverage pre-trained foundation models. A foundation model is a machine learning model trained on large quantities of data that can be adapted to solve a wide variety of downstream tasks. Extending a foundation model to solve a specific problem usually requires less data than training it from scratch and improves the generalizability of the specialist model. Methods will be developed to solve open-world image-level problems requiring natural language input with pre-trained foundation models. Insights from these developments will inspire the construction of models that solve analogous problems in videos. Examples of problems include counting text-specified objects in images and repetitions in videos and answering queries about the area, shape, and structure of objects in a scene. Rather than solving a problem for a particular class, the models developed will allow users to solve the problem for any arbitrary class by providing text input about the class of interest at inference time. Importantly, adapting such open-world models to new classes would require no additional training or data, even if the class were unseen during training. Hence, this work will result in AI systems that are more accessible to the general public, who may not have access to the large quantities of labelled data and compute typically necessary to train class-specific models.Aims & Objectives1. Develop models for image understanding that allow users to ask questions on the image content using natural language.2. Leverage insights and methods from step (1) to develop models with similar capabilities for video understanding that allow users to ask questions on video content using natural language. For instance, a model developed to count objects in images using text in step (1) could inspire a model to count objects in videos using text in step (2).3. Iterate on steps (1) and (2), adding more capabilities. Novelty of the Research MethodologyWhile leveraging pre-trained vision-language foundation models for tasks such as image retrieval, object detection, and instance segmentation has been significantly explored for images, similar developments have been less explored for videos. This is because learning from videos is more complex due to an additional temporal dimension. Furthermore, methods developed in this project will include novel deep learning architectures that are more general and perform better at existing tasks or that solve new problems such as repetition counting in videos using natural language descriptions and answering arbitrary natural language queries about the size, shape, and structure of objects.Alignment to the EPSRC's Strategies & Research AreasThis project relates to the "Artificial Intelligence Technologies" research area.Any Companies or Collaborators Involved?No.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金