Speaker Identification for Business-Card-Type Sensors

Speaker Identification for Business-Card-Type Sensors
复制标题

名片型传感器的说话人识别

DOI:
10.1109/ojcs.2021.3075469
复制
发表时间:
2021
影响因子:
5.9
通讯作者:
Shunpei Yamaguchi;Ritsuko Oshima;J. Oshima;Ryota Shiina;T. Fujihashi;S. Saruwatari;Takashi Watanabe
Shunpei Yamaguchi;Ritsuko Oshima;J. Oshima;Ryota Shiina;T. Fujihashi;S. Saruwatari;Takashi Watanabe
中科院分区:
--
文献类型:
--
作者:
Shunpei Yamaguchi;Ritsuko Oshima;J. Oshima;Ryota Shiina;T. Fujihashi;S. Saruwatari;Takashi Watanabe

文献摘要

相似文献

人的协作对多人活动的性能有很大的影响。通过对说话人信息和语音时序的分析,可以更详细地提取人类协作数据。一些研究通过识别带有名片类型传感器的说话者来提取人类协作数据。然而,由于测量的声压数据中的尖峰、非扬声器传感器中的环境噪声以及每个传感器之间的同步误差,难以以低成本和高精度实现名片型传感器的扬声器识别。本研究提出一种新的声压感测器与说话人辨识演算法,以实现名片式感测器的说话人辨识。该传感器通过采用峰值保持电路和时间同步模块来缓解尖峰和精确的时间同步,以低成本和高精度提取用户的语音。该算法通过去除环境噪声以高精度识别说话人。评估结果表明,该算法准确地识别扬声器在多人活动考虑不同数量的用户,环境噪声,混响条件以及长或短的话语。此外,峰值保持电路能够精确提取语音,传感器之间的同步误差始终在$\pm$30 $\boldsymbol\mu$s以内,即误差可忽略不计。
Human collaboration has a great impact on the performance of multi-person activities. The analysis of speaker information and speech timing can be used to extract human collaboration data in detail. Some studies have extracted human collaboration data by identifying a speaker with business-card-type sensors. However, it is difficult to realize speaker identification for business-card-type sensors at low cost and high accuracy because of spikes in the measured sound pressure data, ambient noise in the non-speaker sensor, and synchronization errors across each sensor. This study proposes a novel sound pressure sensor and speaker identification algorithm to realize speaker identification for business-card-type sensors. The sensor extracts the user's speech at low cost and high accuracy by employing a peak hold circuit and time synchronization module for spike mitigation and precise time synchronization. The algorithm identifies a speaker with high accuracy by removing ambient noise. The evaluations show that the algorithm accurately identifies a speaker in a multi-person activity considering varying numbers of users, environmental noises, and reverberation conditions as well as long or short utterances. In addition, the peak hold circuit enables accurate extraction of speech and the synchronization error between the sensors is always within $\pm$30 $\boldsymbol\mu$s, that is, negligible error.