Acuity: Creating Realistic Digital Twins Through Multi-resolution Pointcloud Processing and Audiovisual Sensor Fusion

Acuity: Creating Realistic Digital Twins Through Multi-resolution Pointcloud Processing and Audiovisual Sensor Fusion
复制标题

DOI:
10.1145/3576842.3582363
复制
发表时间:
2023-05
期刊:
Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation
影响因子:
--
通讯作者:
Jason Wu;Ziqi Wang;Ankur Sarker;M. Srivastava
Jason Wu;Ziqi Wang;Ankur Sarker;M. Srivastava
中科院分区:
其他
文献类型:
--
作者:
Jason Wu;Ziqi Wang;Ankur Sarker;M. Srivastava

文献摘要

相似文献

随着增强和虚拟现实(AR/VR)技术的成熟,希望在具有高忠诚的虚拟场景中以视觉和听觉来代表现实世界中的人们,以制作沉浸式和现实的用户体验。通过化身渲染对象的视觉表示,麦克风阵列被用来定位和分离高质量的主题音频。在视觉域中,挑战仍然存在。然而,为了代表人类的受试者,在音频领域,这种三维数据在计算上是昂贵的。但是,声音源本地化算法可能需要不可接受的时间来提供这些知识,这仍然可能是错误的,尤其是在移动对象的情况下。诸如AR/VR会议之类的应用程序,授权可忽略的系统延迟。在视觉上和极光上的虚拟场景中利用视听传感器融合方法加快声音源分离。没有运行声源的声音,我们的结果表明,敏锐度可以隔离多个受试者的高质量点云,最大延迟为70毫秒,平均吞吐量超过25 fps,而在不到30毫秒内将音频分开。敏锐度守则:https://github.com/nesl/acuity。
As augmented and virtual reality (AR/VR) technology matures, a method is desired to represent real-world persons visually and aurally in a virtual scene with high fidelity to craft an immersive and realistic user experience. Current technologies leverage camera and depth sensors to render visual representations of subjects through avatars, and microphone arrays are employed to localize and separate high-quality subject audio through beamforming. However, challenges remain in both realms. In the visual domain, avatars can only map key features (e.g., pose, expression) to a predetermined model, rendering them incapable of capturing the subjects’ full details. Alternatively, high-resolution point clouds can be utilized to represent human subjects. However, such three-dimensional data is computationally expensive to process. In the realm of audio, sound source separation requires prior knowledge of the subjects’ locations. However, it may take unacceptably long for sound source localization algorithms to provide this knowledge, which can still be error-prone, especially with moving objects. These challenges make it difficult for AR systems to produce real-time, high-fidelity representations of human subjects for applications such as AR/VR conferencing that mandate negligible system latency. We present Acuity, a real-time system capable of creating high-fidelity representations of human subjects in a virtual scene both visually and aurally. Acuity isolates subjects from high-resolution input point clouds. It reduces the processing overhead by performing background subtraction at a coarse resolution, then applying the detected bounding boxes to fine-grained point clouds. Meanwhile, Acuity leverages an audiovisual sensor fusion approach to expedite sound source separation. The estimated object location in the visual domain guides the acoustic pipeline to isolate the subjects’ voices without running sound source localization. Our results demonstrate that Acuity can isolate multiple subjects’ high-quality point clouds with a maximum latency of 70 ms and average throughput of over 25 fps, while separating audio in less than 30 ms. We provide the source code of Acuity at: https://github.com/nesl/Acuity.