课题基金 / 基金详情

Immersive Audio-Visual 3D Scene Reproduction Using a Single 360 Camera

Immersive Audio-Visual 3D Scene Reproduction Using a Single 360 Camera
使用单个 360 度摄像头实现沉浸式视听 3D 场景再现
批准号:
EP/V03538X/1
负责人:
HANSUNG KIM
金额:
$34.08万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The COVID-19 pandemic has changed our lifestyle and caused high demand for remote communication and experience. Many organizations have had to set up remote work systems with video conferencing platforms. However, current video conferencing systems do not meet basic requirements for remote collaboration due to the lack of eye contact, gaze awareness and spatial audio synchronisation. Reproduction of a real space as an audio-visual 3D model allows users to remotely experience real-time interaction in real environments, thus it can be widely utilised in various applications such as healthcare, teleconferencing, education, entertainments, etc. The goal of this project is to develop a simple and practical solution to estimate geometrical structure and acoustic properties of general scenes allowing spatial audio to be adapted to the environment and listener location to give an immersive rendering of the scene to improve user experience.Existing 3D scene reproduction systems have two problems. (i) Audio and vision systems have been researched separately. Computer vision research has mainly focused on improving the visual side of scene reconstruction. In an immersive display, such as a VR system, the experience is not perceived as "realistic" by users if sound is not matched with the visual cues. On the other hand, audio researches have been using only audio sensors to measure acoustic properties without considering the complementary effect with visual sensors. (ii) Current capture and recording systems for 3D scene reproduction require too invasive set up and professional process to be deployed by users in their private places. A LiDAR sensor is expensive and requires long scanning time. Perspective images require large number of photos to cover the whole scene. The objective of this research is to develop an end-to-end audio-visual 3D scene reproduction pipeline using a single shot from a consumer 360 (panoramic) camera. In order to make the system easily accessible by common users in their own private spaces, automatic solution using computer vision and artificial intelligence algorithms should be included in the back-end. A deep neural network (DNN) jointly trained for semantic scene reconstruction and acoustic property prediction for the captured environments will be developed. This process includes inference for invisible regions from the camera. Impulse Responses (IRs) characterising acoustic attributes of an environment allow to reproduce the acoustics of the space with any sound sources. It also allows to extract the original (dry) sound by eliminating acoustic effects from recorded sound so that this source can be re-rendered in new environments with different acoustic effects. A simple and efficient method to estimate acoustic IRs from the captured single 360 photo will be investigated. This semantic scene data is used to provide immersive audio-visual experience to users. Two types of display scenarios will be considered: personalised display system such as a VR headset with headphones and communal display system (e.g., TV or projector) with loudspeakers. Real-time 3D human pose tracking using a single 360 camera will be developed to accurately render 3D audio-visual scene at the locations of users. Delivering binaural sound to listeners using loudspeakers is a challenging task. Audio beam-forming techniques aligned with human-pose tracking for multiple loudspeakers will be investigated in collaboration with the project partners in audio processing. The resulting system would have a significant impact on innovation of VR and multimedia systems, and open up new and interesting applications for their deployment. This award should provide the foundation for the PI to establish and lead a group with a unique research direction which is aligned with national priorities and will address a major long-term research challenge.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
Computer Vision, Imaging and Computer Graphics Theory and Applications - 17th International Joint Conference, VISIGRAPP 2022, Virtual Event, February 6-8, 2022, Revised Selected Papers
计算机视觉、成像和计算机图形理论与应用 - 第 17 届国际联合会议,VISIGRAPP 2022,虚拟活动,2022 年 2 月 6-8 日,修订后的精选论文
DOI: 10.1007/978-3-031-45725-8_4
发表时间: 2023
期刊:
影响因子: --
作者: [Heng Y]
通讯作者: Heng Y
Material Recognition for Immersive Interactions in Virtual/Augmented Reality
虚拟/增强现实中沉浸式交互的材料识别
DOI: 10.1109/vrw58643.2023.00131
发表时间: 2023
期刊:
影响因子: --
作者: [Heng Y]
通讯作者: Heng Y
DOI: 10.5220/0010853200003124
发表时间: 2022
期刊:
影响因子: --
作者: [Yuwen Heng;Yihong Wu;S. Dasmahapatra;Hansung Kim]
通讯作者: Yuwen Heng;Yihong Wu;S. Dasmahapatra;Hansung Kim
DOI: 10.48550/arxiv.2305.03919
发表时间: 2023-05
期刊: ArXiv
影响因子: --
作者: [Yuwen Heng;S. Dasmahapatra;Hansung Kim]
通讯作者: Yuwen Heng;S. Dasmahapatra;Hansung Kim
9
    海外基金