课题基金 / 基金详情

Audio-Visual Egocentric Video Understanding

Audio-Visual Egocentric Video Understanding
视听以自我为中心的视频理解
批准号:
2615061
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
By nature, human learning is multi-modal. We combine information from multiple inputs via our senses, such as touch, sound, sight, and gain a better understanding of the world. Commonly, we will combine these modalities in order to learn how to do complete tasks. For example, consider when we try to learn a musical instrument, we will typically utilise both our sight and hearing to understand how different keys on a piano, or frets on a guitar, will ring out different sounds and therefore create music. Furthermore, there are instances when one modality can assist in understanding the continuation of an action, despite another modality shifting, such as when watching a chef fry something in a frying pan in cooking video; if the camera shifts and no longer visually shows what is in the pan, we can still use the sizzling sound to understand that the frying is still taking place, despite a visual shift. However, in the context of deep learning, auditory data linked to the video stream is a commonly underutilised resource, and potential increases in performance from integrating this audio data is often left neglected. Therefore, it seems logical to attempt to design and optimize audio-visual models for video understanding tasks to both: better model multi-modal human learning and also to improve performance over uni-modal solutions.However, this is no trivial challenge, as it is not simply a case of optimizing each modality separately and then combining them together. There are multiple nuances and considerations with combining the modalities, which will constantly change between video understanding tasks, datasets, architectures, and other aspects of deep learning. This includes: how do we fuse the audio and video streams? At what point in the model do we fuse them? Once they are fused, how do we allow the modalities to communicate between each other? In our work, we seek to answer these questions, investigating a wide spectrum of audio-visual action recognition methods with the aim of improving accuracy results within the domain of action recognition, whilst developing and training models on the large-scale egocentric dataset Epic-Kitchens. This project relates to the image and vision computing research area within EPSRC, with its most obvious real-world application being applied to robot learning, which is further assisted by the egocentric (first-person) nature of the video data we use. However, this work can apply to any real-world applications that involve computer vision and is not necessarily restricted to robotics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
  • 批准号:
    60373031
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2003
  • 负责人:
    陈越
  • 依托单位: