Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings

Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings
复制标题

DOI:
10.1109/slt48900.2021.9383578
复制
发表时间:
2020-07
期刊:
2021 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Bowen Shi;Shane Settle;Karen Livescu
Bowen Shi;Shane Settle;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
Bowen Shi;Shane Settle;Karen Livescu

文献摘要

被引文献

相似文献

分段模型是序列预测模型,其中假设的分数基于帧的整个可变长度段。我们认为分段模型的全字(“声字”)语音识别,使用向量嵌入段定义的特征向量。这样的模型在计算上是具有挑战性的,因为路径的数量与词汇量大小成比例,这可能比使用子字单元(如音素)时大几个数量级。我们描述了一种有效的方法,端到端的全词分段模型,与前向-后向和Viterbi解码上的GPU和一个简单的段评分功能,降低了空间的复杂性。此外,我们还研究了通过联合训练的声学词嵌入(AWE)和声学接地词嵌入(AGWE)的书面词标签进行预训练的使用。我们发现,字错误率可以通过使用AWE预训练声学段表示来大幅降低,并且可以通过使用AGWE预训练字预测层来获得额外的(较小的)增益。我们的最终模型比之前的A2 W模型有所改进。
Segmental models are sequence prediction models in which scores of hypotheses are based on entire variable-length segments of frames. We consider segmental models for whole-word ("acoustic-to-word") speech recognition, with the feature vectors defined using vector embeddings of segments. Such models are computationally challenging as the number of paths is proportional to the vocabulary size, which can be orders of magnitude larger than when using subword units like phones. We describe an efficient approach for end-to-end whole-word segmental models, with forward-backward and Viterbi decoding performed on a GPU and a simple segment scoring function that reduces space complexity. In addition, we investigate the use of pre-training via jointly trained acoustic word embeddings (AWEs) and acoustically grounded word embeddings (AGWEs) of written word labels. We find that word error rate can be reduced by a large margin by pre-training the acoustic segment representation with AWEs, and additional (smaller) gains can be obtained by pre-training the word prediction layer with AGWEs. Our final models improve over prior A2W models.