An Audio-Visual System for Object-Based Audio: From Recording to Listening

An Audio-Visual System for Object-Based Audio: From Recording to Listening
复制标题

DOI:
10.1109/tmm.2018.2794780
复制
发表时间:
2018-01
影响因子:
7.3
通讯作者:
Philip Coleman;Andreas Franck;J. Francombe;Qingju Liu;Teofilo de Campos;R. Hughes;D. Menzies;M. S. Gálvez;Yan Tang;James Woodcock;P. Jackson;F. Melchior;C. Pike;F. Fazi;T. Cox;A. Hilton
Philip Coleman;Andreas Franck;J. Francombe;Qingju Liu;Teofilo de Campos;R. Hughes;D. Menzies;M. S. Gálvez;Yan Tang;James Woodcock;P. Jackson;F. Melchior;C. Pike;F. Fazi;T. Cox;A. Hilton
中科院分区:
计算机科学1区
文献类型:
--
作者:
Philip Coleman;Andreas Franck;J. Francombe;Qingju Liu;Teofilo de Campos;R. Hughes;D. Menzies;M. S. Gálvez;Yan Tang;James Woodcock;P. Jackson;F. Melchior;C. Pike;F. Fazi;T. Cox;A. Hilton

文献摘要

被引文献

相似文献

基于对象的音频是音频内容的新兴表示,其中内容以再现格式不可知的方式表示,并且因此产生一次以供在许多不同种类的设备上消费。这为沉浸式、个性化和交互式聆听体验提供了新的机会。本文介绍了一种端到端的基于对象的空间音频管道,从录音到收听。提出了一个高层次的系统架构,其中包括新颖的视听接口,以支持基于对象的捕获和跟踪渲染,并纳入了一个建议的组件的对象化,即直接记录到一个基于对象的形式的内容。基于文本的可扩展元数据支持系统组件之间的通信。提出了一种开放的对象绘制体系结构。该系统的能力进行了评估,在两个部分。首先,使用客观双耳定位模型来评估从两个移动讲话者自动估计的元数据的录音机跟踪再现。其次,基于对象的场景捕获与使用盲源分离(两个谈话者之间的混音)和波束成形(混音爵士乐队的录音)提取的音频进行评估与感知动机的客观和主观的实验。这些实验表明,该系统的新组件增加了超越最先进的能力。最后,我们讨论了基于对象的音频工作流程的挑战和未来的前景。
Object-based audio is an emerging representation for audio content, where content is represented in a reproduction-format-agnostic way and, thus, produced once for consumption on many different kinds of devices. This affords new opportunities for immersive, personalized, and interactive listening experiences. This paper introduces an end-to-end object-based spatial audio pipeline, from sound recording to listening. A high-level system architecture is proposed, which includes novel audio-visual interfaces to support object-based capture and listener-tracked rendering, and incorporates a proposed component for objectification, that is, recording content directly into an object-based form. Text-based and extensible metadata enable communication between the system components. An open architecture for object rendering is also proposed. The system's capabilities are evaluated in two parts. First, listener-tracked reproduction of metadata automatically estimated from two moving talkers is evaluated using an objective binaural localization model. Second, object-based scene capture with audio extracted using blind source separation (to remix between two talkers) and beamforming (to remix a recording of a jazz group) is evaluated with perceptually motivated objective and subjective experiments. These experiments demonstrate that the novel components of the system add capabilities beyond the state of the art. Finally, we discuss challenges and future perspectives for object-based audio workflows.