Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR

Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR
复制标题

多说话者端到端 ASR 的扩展图时间分类

DOI:
10.48550/arxiv.2203.00232
复制
发表时间:
2022
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Jonathan Le Roux
Jonathan Le Roux
中科院分区:
--
文献类型:
--
作者:
Xuankai Chang;Niko Moritz;Takaaki Hori;Shinji Watanabe;Jonathan Le Roux

文献摘要

被引文献

相似文献

基于图的时间分类(GTC)是连接主义时间分类损失的一种广义形式,近年来被提出用于改进基于图的自动语音识别系统。例如,GTC首先用于将伪标签序列的N-best列表编码为半监督学习的图。在本文中,我们提出了对GTC的扩展,通过神经网络对标签和标签转换的后置进行建模,这可以应用于更广泛的任务。作为一个示例应用,我们将扩展的GTC (GTC-e)用于多说话人语音识别任务。多说话人语音的转录和说话人信息用图表示,其中说话人信息与转换相关联,ASR输出与节点相关联。使用GTC-e,多扬声器ASR建模变得与单扬声器ASR建模非常相似,因为多个扬声器的令牌被识别为按时间顺序合并的单个序列。为了进行评估,我们在来自librisspeech的模拟多演讲者语音数据集上进行了实验,获得了接近经典基准的有希望的结果。
Graph-based temporal classification (GTC), a generalized form of the connectionist temporal classification loss, was recently proposed to improve automatic speech recognition (ASR) systems using graph-based supervision. For example, GTC was first used to encode an N-best list of pseudo-label sequences into a graph for semi-supervised learning. In this paper, we propose an extension of GTC to model the posteriors of both labels and label transitions by a neural network, which can be applied to a wider range of tasks. As an example application, we use the extended GTC (GTC-e) for the multi-speaker speech recognition task. The transcriptions and speaker information of multi-speaker speech are represented by a graph, where the speaker information is associated with the transitions and ASR outputs with the nodes. Using GTC-e, multi-speaker ASR modelling becomes very similar to single-speaker ASR modeling, in that tokens by multiple speakers are recognized as a single merged sequence in chronological order. For evaluation, we perform experiments on a simulated multi-speaker speech dataset derived from LibriSpeech, obtaining promising results with performance close to classical benchmarks for the task.