Audio-Visual Egocentric Video Understanding
Audio-Visual Egocentric Video Understanding
批准号:
2615061
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
从本质上讲,人类的学习是多模式的。我们通过我们的感官,如触觉、声音、视觉,将来自多种输入的信息组合在一起,从而更好地理解世界。通常,我们会结合这些模式来学习如何完成任务。例如,当我们尝试学习一种乐器时,我们通常会利用我们的视觉和听觉来理解钢琴上不同的键或吉他上的拨弦如何发出不同的声音,从而创造出音乐。此外,在某些情况下,一种情态可以帮助理解动作的继续,尽管另一种情态发生了变化,比如在烹饪视频中看到厨师在煎锅里煎东西;如果镜头换了,看不到锅里有什么,我们仍然可以用滋滋的声音来理解煎炸仍然在进行,尽管视觉上发生了变化。然而,在深度学习的背景下,与视频流相关联的听觉数据通常是一种未被充分利用的资源,并且通过集成音频数据来提高性能的潜力往往被忽视。因此,尝试为视频理解任务设计和优化视听模型似乎是合乎逻辑的:更好地模拟多模态人类学习,并提高单模态解决方案的性能。然而,这不是一个简单的挑战,因为它不是简单地分别优化每种模式,然后将它们组合在一起的情况。结合模式有许多细微差别和考虑因素,这些模式将在视频理解任务、数据集、架构和深度学习的其他方面不断变化。这包括:我们如何融合音频和视频流?我们在模型的哪个点融合它们?一旦它们融合在一起,我们如何让这些模式相互交流呢?在我们的工作中,我们试图回答这些问题,研究广泛的视听动作识别方法,目的是提高动作识别领域内的准确性结果,同时在大规模自我中心数据集Epic-Kitchens上开发和训练模型。该项目涉及EPSRC的图像和视觉计算研究领域,其最明显的现实应用是应用于机器人学习,这进一步得益于我们使用的视频数据的自我中心(第一人称)性质。然而,这项工作可以应用于任何涉及计算机视觉的现实世界的应用,而不一定局限于机器人。
英文摘要
By nature, human learning is multi-modal. We combine information from multiple inputs via our senses, such as touch, sound, sight, and gain a better understanding of the world. Commonly, we will combine these modalities in order to learn how to do complete tasks. For example, consider when we try to learn a musical instrument, we will typically utilise both our sight and hearing to understand how different keys on a piano, or frets on a guitar, will ring out different sounds and therefore create music. Furthermore, there are instances when one modality can assist in understanding the continuation of an action, despite another modality shifting, such as when watching a chef fry something in a frying pan in cooking video; if the camera shifts and no longer visually shows what is in the pan, we can still use the sizzling sound to understand that the frying is still taking place, despite a visual shift. However, in the context of deep learning, auditory data linked to the video stream is a commonly underutilised resource, and potential increases in performance from integrating this audio data is often left neglected. Therefore, it seems logical to attempt to design and optimize audio-visual models for video understanding tasks to both: better model multi-modal human learning and also to improve performance over uni-modal solutions.However, this is no trivial challenge, as it is not simply a case of optimizing each modality separately and then combining them together. There are multiple nuances and considerations with combining the modalities, which will constantly change between video understanding tasks, datasets, architectures, and other aspects of deep learning. This includes: how do we fuse the audio and video streams? At what point in the model do we fuse them? Once they are fused, how do we allow the modalities to communicate between each other? In our work, we seek to answer these questions, investigating a wide spectrum of audio-visual action recognition methods with the aim of improving accuracy results within the domain of action recognition, whilst developing and training models on the large-scale egocentric dataset Epic-Kitchens. This project relates to the image and vision computing research area within EPSRC, with its most obvious real-world application being applied to robot learning, which is further assisted by the egocentric (first-person) nature of the video data we use. However, this work can apply to any real-world applications that involve computer vision and is not necessarily restricted to robotics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
-
批准号:60373031
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2003
-
负责人:陈越
-
依托单位: