课题基金 / 基金详情

Challenges in Immersive Audio Technology

Challenges in Immersive Audio Technology
沉浸式音频技术的挑战
批准号:
EP/X032981/1
负责人:
Zoran Cvetkovic
金额:
$121.51万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --

项目摘要

项目成果

Zoran Cvetkovic的其他基金

相似基金

相关文献

中文摘要
翻译
沉浸式技术不仅将改变我们的交流和娱乐体验方式,还将改变我们对物理世界的体验,从商店到博物馆,从汽车到教室。这种转变主要是由视觉技术的前所未有的进步所推动的,这些技术使用户能够进入另一种视觉现实。然而,在音频领域,需要克服长期存在的基本挑战,以实现引人注目的沉浸式体验,其中一组听众只需走进一个场景,就可以感受到另一个现实,享受无缝的共享体验,而无需耳机,头部跟踪,个性化或校准。第一个关键挑战是向多个听众提供沉浸式音频体验。音频技术的最新进展开始成功地产生高质量的沉浸式音频体验。然而,这些在实践中被限制为个体收听者,其中适当的信号经由耳机或经由基于使用串扰消除或波束成形的适度数量的扬声器的系统呈现。在技术上有效地将“3D声音”传递给多个听众方面仍然存在根本性的挑战,无论是在家庭环境中的小数量(2-5个)、在博物馆、画廊和其他公共场所(5-20个)还是在电影院和剧院剧院(20-100个)。原则上,可以使用基于物理的方法(诸如波场合成或高阶立体混响)来生成共享的听觉体验,但是即使是适度大小的最佳点也需要数量惊人的通道。CIAT旨在通过开发一个原则性的可扩展和可重新配置的框架来捕获和再现感知相关的信息,从而在实际可行的系统可实现的沉浸式音频体验的质量方面取得进步,从而改变现有技术。第二个关键挑战是将听众传送到替代现实所需的环境声学的实时计算,允许它们与环境和其中的声源交互。这与沉浸式音频内容被合成而不是被记录的应用以及通常基于对象的音频有关。声事件的声场包括直接波前,随后是早期和高阶反射。一个令人信服的经验,被运送到环境中的事件发生,需要渲染这些反射,这不能都在真实的时间计算。在真实感至关重要的应用中,例如延展实境(XR)和某种程度上的游戏,环境的脉冲响应通常仅在几个位置处计算,具有对反射数量和到达方向的预设限制,然后与源声音卷积以实现所谓的高质量混响。尽管如此,冲激响应和卷积的计算可能需要GPU实现,并在质量和复杂性之间以及CPU和GPU计算之间进行仔细的实际平衡。CIAT旨在实现环境建模的范式转变,从而实现真实的数字高效无缝高质量环境模拟。通过应对这些挑战,CIAT将为新兴XR应用创建和交付共享的交互式沉浸式音频体验,同时在传统媒体的沉浸式音频质量方面取得进步。特别是,高质量环境声学的高效实时合成对于XR和基于对象的音频(包括流媒体和广播)都是必不可少的。向多个听众提供3D音景在传统应用中也是一个尚未解决的主要问题,包括广播,电影,音乐活动和视听装置。
英文摘要
Immersive technologies will transform not only how we communicate and experience entertainment, but also our experience of the physical world, from shops to museums, cars to classrooms. This transformation has been driven primarily by an unprecedented progress in visual technologies, which enable transporting users to an alternate visual reality. In the domain of audio, there are however long-standing fundamental challenges that need to be overcome to enable striking immersive experiences in which a group of listeners can just walk into a scene and feel transported to an alternate reality to enjoy a seamless shared experience without the need for headphones, head-tracking, personalisation or calibration.The first key challenge is the delivery of immersive audio experiences to multiple listeners. Recent advances in audio technology are beginning to succeed in generating high quality immersive audio experiences. However, these are restricted in practice to individual listeners, with appropriate signals presented either via headphones, or via systems based on a modest number of loudspeakers using either cross-talk cancellation or beamforming. There remains a fundamental challenge in the technologically efficient delivery of "3D sound" to multiple listeners, either in small numbers (2-5) in a home environment, in museums, galleries and other public spaces (5-20) or in cinema and theatre auditoria (20-100). In principle, shared auditory experiences can be generated using physics-based methods such as wavefield synthesis or higher order ambisonics, but a sweet spot of even a modest size requires a prohibitive number of channels. CIAT aims to transform state of the art by developing a principled scalable and reconfigurable framework for capturing and reproducing only perceptually relevant information, thus leading to a step advance in the quality of immersive audio experiences achievable by practically viable systems.The second key challenge is the real-time computation of environment acoustics needed to transport listeners to alternate reality, allowing them to interact with the environment and sound sources in it. This is pertinent to applications where immersive audio content is synthesised rather than recorded and to object-based audio in general. The sound field of an acoustic event consists of direct wavefront, followed by early and higher-order reflections. A convincing experience of being transported to the environment where the event takes place requires the rendering of these reflections, which cannot all be computed in real time. In applications where the sense of realism is critical, e.g. extended reality (XR) and to some extent gaming, impulse responses of the environment are typically computed only at several locations, with preset limits on the number reflections and directions of arrival, and then convolved with source sounds to achieve what is referred to as high-quality reverberation. Still, the computation of impulse responses and convolution may require GPU implementation and careful hands-on balancing between quality and complexity, and between CPU and GPU computation. CIAT aims to deliver a paradigm shift in environment modelling that will enable numerically efficient seamless high quality environment simulation in real time.By addressing these challenges, CIAT will enable creation and delivery of shared interactive immersive audio experiences for emerging XR applications, whilst making a step advance in the quality of immersive audio in traditional media. In particular, efficient real-time synthesis of high quality environment acoustics is essential for both XR and object-based audio in general, including streaming and broadcasting. Delivery of 3D soundscapes to multiple listeners is a major unresolved problem in traditional applications too, including broadcasting, cinema, music events, and audio-visual installations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SpeechWave
  • 批准号:
    EP/R012067/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $93.54万
  • 财政年份:
    2018
  • 负责人:
    Zoran Cvetkovic
  • 依托单位:
Visits to University of California, Berkeley, Stanford University, and SRI International
  • 批准号:
    EP/K034626/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $2.68万
  • 财政年份:
    2013
  • 负责人:
    Zoran Cvetkovic
  • 依托单位:
Perceptual Sound Field Reconstruction and Coherent Emulation
  • 批准号:
    EP/F001142/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $49.67万
  • 财政年份:
    2008
  • 负责人:
    Zoran Cvetkovic
  • 依托单位:
Robust Syllable Recognition in the Acousic-Waveform Domain
  • 批准号:
    EP/D053005/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $26.44万
  • 财政年份:
    2006
  • 负责人:
    Zoran Cvetkovic
  • 依托单位:
海外基金