Panoptic Reconstruction of Immersive Virtual Soundscapes Using Human-Scale Panoramic Imagery with Visual Recognition

Panoptic Reconstruction of Immersive Virtual Soundscapes Using Human-Scale Panoramic Imagery with Visual Recognition
复制标题

DOI:
10.21785/icad2021.043
复制
发表时间:
2021-06
期刊:
Proceedings of the 26th International Conference on Auditory Display (ICAD 2021)
影响因子:
--
通讯作者:
Mincong Huang;Samuel Chabot;J. Braasch
Mincong Huang;Samuel Chabot;J. Braasch
中科院分区:
其他
文献类型:
--
作者:
Mincong Huang;Samuel Chabot;J. Braasch

文献摘要

相似文献

这项工作位于Rensselaer的协作研究增强沉浸式虚拟环境实验室(CRAIVE-Lab),使用全景图像数据集进行空间音频显示。为以房间为中心的沉浸式虚拟现实设施开发了一种系统,以使用用于语义分割和对象检测的预先训练的神经网络模型逐段地分析全景图像,从而生成具有各自空间位置的音频对象。然后,这些音频对象与一系列合成和录制的音频数据集进行映射,并作为虚拟声源填充在空间音频环境中。所产生的视听结果随后使用该设施的人体尺寸全景显示器以及用于波场合成(WFS)的128声道扬声器阵列来显示。性能评估表明实时增强的有效性,具有在动态沉浸式虚拟环境中大规模扩展和快速部署的潜力。
This work, situated at Rensselaer’s Collaborative-Research Augmented Immersive Virtual Environment Laboratory (CRAIVE-Lab), uses panoramic image datasets for spatial audio display. A system is developed for the room-centered immersive virtual reality facility to analyze panoramic images on a segment-by-segment basis, using pre-trained neural network models for semantic segmentation and object detection, thereby generating audio objects with respective spatial locations. These audio objects are then mapped with a series of synthetic and recorded audio datasets and populated within a spatial audio environment as virtual sound sources. The resulting audiovisual outcomes are then displayed using the facility’s human-scale panoramic display, as well as the 128-channel loudspeaker array for wave field synthesis (WFS). Performance evaluation indicates effectiveness for real-time enhancements, with potentials for large-scale expansion and rapid deployment in dynamic immersive virtual environments.