Perceptual Sound Field Reconstruction and Coherent Emulation
Perceptual Sound Field Reconstruction and Coherent Emulation
批准号:
EP/F001142/1
负责人:
Zoran Cvetkovic
金额:
$49.67万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --
中文摘要
该项目涉及开发一种新的5- 10声道音频技术,该技术将在以下方面改进现有技术:(a)真实感,(b)听觉视角的准确性和稳定性,(c)最佳点的大小,以及(d)包络体验。由于新技术旨在创造360度的听觉视角,因此再现将通过位于正多边形顶点的扬声器进行。每个扬声器将由两个组件组成,一个将直接向听众辐射声场,另一个将通过引入额外的散射来再现漫射声场。下面列出的特定任务的目标是找到捕获声场线索的最佳方法,并使用拟议的重放系统以一种能够提供最令人信服的原始或期望声场错觉的方式呈现它们。(i)将研究拟议的回放系统的最佳麦克风阵列。考虑的阵列将由放置在正多边形顶点的水平面上的麦克风组成,麦克风的数量等于扬声器的数量。对于每个阵列,将考虑不同的直径,从接近重合到略高于最佳值的范围,以及不同的传声器指向性模式。这些研究将对扬声器配置的几个直径进行重复,以调查最佳阵列直径是否取决于扬声器布局的大小,如果是这样,则表征这种依赖关系。最佳传声器指向性模式和阵列直径之间可能存在的依赖关系也将被研究和表征。在关键的听力测试中,将根据上述标准(a)至(d)对阵列进行评估。实验将以模拟为指导,模拟将提供对听区内产生的过渡段和过渡段信号的初步客观评估。同时,本文还将研究由该技术产生的声场的数学模型,这将为最佳麦克风阵列设计提供一些额外的见解。(ii)将有系统地调查取消相声回放的影响。将首先使用现有的串音消除算法,如有必要,将开发在一系列收听环境中具有数字效率和有效性的新算法。然后,将研究具有串音消除回放的最佳麦克风阵列,即(i)中描述的工作将重复用于具有串音消除的复制。最后,比较了有无串音消除的最优系统。将研究直接/扩散声场分离的算法。当仪器的数量不超过麦克风的数量时,可以使用多通道均衡技术来找到干源信号,然后将其与房间脉冲响应的直接/混响部分进行卷积,分别获得直接/漫射声场分量。然而,由于脉冲响应过长,音频中的多通道均衡特别具有挑战性,我们将开发用于音频应用的多通道均衡的数值高效算法。然后,我们将研究不受声源数量限制的直接/扩散声场分解的心理声学近似。(四)根据标准(a)至(d),将在关键的听力测试中系统地研究用于获取直接声场线索的接近一致的定向麦克风阵列和用于获取漫射声场线索的基于全向或双向麦克风的宽间隔阵列的组合。此方法将与(i)- (iii)中描述的方法进行比较,其中两个声场组件使用相同的阵列。
英文摘要
The project is concerned with the development of a new 5--10 channel audio technology which would improve over existing ones in terms of (a) realism, (b) accuracy and stability of the auditory perspective, (c) size of the sweet spot, and (d) the envelopment experience. Since the new technology aims to create a 360 degrees auditory perspective, the reproduction will take place over speakers positioned at vertices of a regular polygon. Each speaker will consist of two components, one which will radiate the direct sound field toward a listener, and another which will reproduce diffuse sound field by introducing additional scattering. The goal of the particular tasks, listed below, is to find optimal ways to capture sound field cues and render them using the proposed playback system in a manner which would provide the most convincing illusion of the original or desired sound field.(i) Optimal microphone arrays for the proposed play-back system will be investigated. Arrays considered will consist of microphones placed in the horizontal plane at the vertices of a regular polygon, with the number of microphones equal to the number of speakers. For each array, different diameters, in the range from near coincident up to somewhat beyond the optimal value, and different microphone directivity patterns will be considered. These studies will be repreated for a few diameters of the speaker configuration to investigate if the optimal array diameter depends on the size of the speaker lay-out, and if so to characterize that dependence. Possible dependencies between the optimal microphone directivity patterns and array diameters will be also investigated and characterized. Arrays will be evaluated in critical listening tests according to criteria (a)--(d) stated in the above. Experiments will be guided by simulations which would provide initial objective assessment of ITD and ILD cues generated within the listening area. In parallel, mathematical models of sound fields generated by the proposed technology will be investigated, which could provide some additional insight into the optimal microphone array design. (ii) The impact of play-back with cross-talk cancellation will be be systematically investigated. Existing cross-talk cancellation algorithms will be first used, and if necessary, new algorithms which are numerically efficient and effective in a range of listening environments will be developed. Then optimal microphone arrays for play back with cross-talk cancellation will be investigated, i.e. the work described under (i) will be repeated for reproduction with cross-talk cancellation. Finally, optimal systems with and without cross-talk cancellation will be compared.(iii) Algorithms for direct/diffuse sound field separation will be studied. When the number of instruments does not exceed the number of microphones, multichannel equalization techniques can be used to find dry source signals, which can then be convolved with direct/reverberant parts of room impulse responses to obtain direct/diffuse sound field components, respectively. Multichannel equalization in audio is, however, particularly challenging owing to excessively long impulse responses, and we will develop numerically efficient algorithms for multichannel equalization for audio applications. Then we will study psychoacoustic approximation to direct/diffuse sound field decomposition with no restriction on the number of sources. (iv) Combinations of near-coincident directional microphone arrays, for acquiring direct sound field cues, and widely spaced arrays based on omni-directional or bi-directional microphones, for acquiring diffuse sound field cues, will be systematically investigated in critical listening tests according to criteria (a)--(d). This approach will be evaluated in comparison with the approach described in (i)--(iii) where the same array is used for both sound field components.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Scattering Delay Network: an interactive reverberator for computer games
散射延迟网络:计算机游戏的交互式混响器
DOI:
--
发表时间:
期刊:
Audio for Games
影响因子:
--
作者:
[Enzo De Sena (Author)]
通讯作者:
Enzo De Sena (Author)
Perceptual evaluation of a circularly symmetric microphone array for panoramic recording of audio
用于全景音频录制的圆形对称麦克风阵列的感知评估
DOI:
--
发表时间:
2010
期刊:
影响因子:
--
作者:
[E. D. Sena, H. Hacıhabiboğlu, Z. Cvetković]
通讯作者:
Z. Cvetković
DOI:
10.1109/taslp.2015.2438547
发表时间:
2015-02
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[E. D. Sena;H. Hacıhabiboğlu;Z. Cvetković;J. Smith]
通讯作者:
E. D. Sena;H. Hacıhabiboğlu;Z. Cvetković;J. Smith
DOI:
10.1109/taslp.2020.2975419
发表时间:
2020
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[De Sena E]
通讯作者:
De Sena E
Analysis and Design of Multichannel Systems for Perceptual Sound Field Reconstruction
感知声场重建多通道系统分析与设计
DOI:
10.1109/tasl.2013.2260152
发表时间:
2013
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[De Sena E]
通讯作者:
De Sena E
共 6 条
Challenges in Immersive Audio Technology
-
批准号:EP/X032981/1
-
项目类别:Research Grant
-
资助金额:$121.51万
-
财政年份:2024
-
负责人:Zoran Cvetkovic
-
依托单位:
SpeechWave
-
批准号:EP/R012067/1
-
项目类别:Research Grant
-
资助金额:$93.54万
-
财政年份:2018
-
负责人:Zoran Cvetkovic
-
依托单位:
Visits to University of California, Berkeley, Stanford University, and SRI International
-
批准号:EP/K034626/1
-
项目类别:Research Grant
-
资助金额:$2.68万
-
财政年份:2013
-
负责人:Zoran Cvetkovic
-
依托单位:
Robust Syllable Recognition in the Acousic-Waveform Domain
-
批准号:EP/D053005/1
-
项目类别:Research Grant
-
资助金额:$26.44万
-
财政年份:2006
-
负责人:Zoran Cvetkovic
-
依托单位:
海外基金