Spatial Audio Empowered Smart speakers with Xblock - A Pose-Adaptive Crosstalk Cancellation Algorithm for Free-moving Users

Spatial Audio Empowered Smart speakers with Xblock - A Pose-Adaptive Crosstalk Cancellation Algorithm for Free-moving Users
复制标题

DOI:
10.1145/3576914.3589563
复制
发表时间:
2023-05
期刊:
Proceedings of Cyber-Physical Systems and Internet of Things Week 2023
影响因子:
--
通讯作者:
F. Liu;Anish Narsipur;Andrew Kemeklis;Lucy Song;R. Likamwa
F. Liu;Anish Narsipur;Andrew Kemeklis;Lucy Song;R. Likamwa
中科院分区:
其他
文献类型:
--
作者:
F. Liu;Anish Narsipur;Andrew Kemeklis;Lucy Song;R. Likamwa

文献摘要

相似文献

智能物联网扬声器虽然通过网络连接,但目前仅产生直接来自各个设备的声音。我们设想未来智能扬声器协作产生空间音频结构,能够感知地将声音放置在物理空间的一系列位置。这可以在家庭、办公室和公共空间中提供音频提示,这些提示可以灵活地连接到各个位置。空间化音频的感知依赖于双耳线索,尤其是用户左耳和右耳入射声音的时间差和电平差。由于听觉串扰,传统立体声扬声器在播放双耳音频时无法为用户创建空间感知,因为每只耳朵都会听到两个扬声器输出的组合。我们提出了 Xblock,一种新颖的时域姿势自适应串扰消除技术,该技术利用用户头部姿势和扬声器位置的知识在一对扬声器上创建空间音频感知。我们构建了一个由Xblock赋能的原型智能音箱物联网系统,通过信号分析探索Xblock的有效性,并讨论未来的感知用户研究和未来的工作。
Smart IoT Speakers, while connected over a network, currently only produce sounds that come directly from the individual devices. We envision a future where smart speakers collaboratively produce a fabric of spatial audio, capable of perceptually placing sound in a range of locations in physical space. This could provide audio cues in homes, offices and public spaces that are flexibly linked to various positions. The perception of spatialized audio relies on binaural cues, especially the time difference and the level difference of incident sound at a user’s left and right ears. Traditional stereo speakers cannot create the spatialization perception for a user when playing binaural audio due to auditory crosstalk, as each ear hears a combination of both speaker outputs. We present Xblock, a novel time-domain pose-adaptive crosstalk cancellation technique that creates a spatial audio perception over a pair of speakers using knowledge of the user’s head pose and speaker positions. We build a prototype smart speaker IoT system empowered by Xblock, explore the effectiveness of Xblock through signal analysis, and discuss future perceptual user studies and future work.