A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network

A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network
复制标题

用于特设声学传感器网络的会议转录系统

DOI:
10.48550/arxiv.2205.00944
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
Reinhold Haeb
Reinhold Haeb
中科院分区:
--
文献类型:
--
作者:
Tobias Gburrek;Christoph Boeddeker;Thilo von Neumann;Tobias Cord;Joerg Schmalenstroeer;Reinhold Haeb

文献摘要

被引文献

相似文献

我们提出了一个系统,该系统记录了一个典型会议场景的对话,该对话由一组位于未知位置的初始不同步麦克风阵列捕获。它由信号同步子系统组成,包括采样率和采样时间偏移估计、基于扬声器和麦克风阵列位置估计的拨号化、多通道语音增强和自动语音识别。利用估计的偏振化信息初始化空间混合模型,用于估计源分离的波束形成系数。仿真结果表明,与单个紧凑麦克风阵列相比,多个分布式麦克风阵列同步组合可以提高语音识别精度。此外,所提出的空间混合模型的知情初始化比随机初始化提供了明显的性能优势。
We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It consists of subsystems for signal synchronization, including both sampling rate and sampling time offset estimation, diarization based on speaker and microphone array position estimation, multi-channel speech enhancement, and automatic speech recognition. With the estimated diarization information, a spatial mixture model is initialized that is used to estimate beamformer coefficients for source separation. Simulations show that the speech recognition accuracy can be improved by synchronizing and combining multiple distributed microphone arrays compared to a single compact microphone array. Furthermore, the proposed informed initialization of the spatial mixture model delivers a clear performance advantage over random initialization.