MAVL: Multiresolution Analysis of Voice Localization

MAVL: Multiresolution Analysis of Voice Localization
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

智能扬声器根据用户的声音定位用户的能力为许多新应用打开了大门。在本文中,我们提出了一种新的系统,MAVL,定位人类的声音。它包括三个主要部分:(i)我们首先开发了一种新的多分辨率分析来估计来自多个传播路径的时变低频相干语音信号的AoA;(ii)然后我们通过发射声学信号和开发改进的3D MUSIC算法来自动估计房间结构;(iii)我们最后使用估计的AoA和房间结构来重新追踪路径以定位语音。我们实现了一个原型系统,使用一个单一的扬声器和一个均匀的圆形麦克风阵列。我们的研究结果表明,它实现了1.49o和3.33o的前两个AoA估计的中位误差,并实现了0.31米的视线(LoS)和0.47米的非视线(NLoS)的情况下,中位定位误差。
The ability for a smart speaker to localize a user based on his/her voice opens the door to many new applications. In this paper, we present a novel system, MAVL, to localize human voice. It consists of three major components: (i) We first develop a novel multi-resolution analysis to estimate the AoA of time-varying low-frequency coherent voice signals coming from multiple propagation paths; (ii) We then automatically estimate the room structure by emitting acoustic signals and developing an improved 3D MUSIC algorithm; (iii) We finally re-trace the paths using the estimated AoA and room structure to localize the voice. We implement a prototype system using a single speaker and a uniform circular microphone array. Our results show that it achieves median errors of 1.49o and 3.33o for the top two AoAs estimation and achieves median localization errors of 0.31m in line-of-sight (LoS) cases and 0.47m in non-line-of-sight (NLoS) cases.