Phonetic Error Analysis Beyond Phone Error Rate

Phonetic Error Analysis Beyond Phone Error Rate
复制标题

DOI:
10.1109/taslp.2023.3313417
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Erfan Loweimi;Andrea Carmantini;Peter Bell;Steve Renals;Z. Cvetkovic
Erfan Loweimi;Andrea Carmantini;Peter Bell;Steve Renals;Z. Cvetkovic
中科院分区:
其他
文献类型:
--
作者:
Erfan Loweimi;Andrea Carmantini;Peter Bell;Steve Renals;Z. Cvetkovic

文献摘要

相似文献

在这篇文章中,我们分析了基于TIMIT的电话识别系统的性能超出了整体电话错误率(PER)度量。我们考虑了三个广义的语音类(BPC):{塞擦音,双元音,摩擦音,鼻音,爆破音,半元音,元音,沉默},{辅音,元音,沉默}和{浊音,清音,沉默},并计算每个语音类的贡献的替代,删除,插入和PER。此外,对于每个BPC,我们研究了以下内容:训练期间PER的演变,噪声的影响(NTIMIT),不同频谱子带(1,2,4和8 kHz)的重要性,双向与单向顺序建模的有用性,WSJ的迁移学习和通过单音的正则化。此外,我们为每个BPC构建了一个混淆矩阵,并通过在声学模型的输入(声学特征)和输出(logits)水平上降维到2D来分析混淆。我们还比较了基于BLSTM的混合基线系统与基于GMM-HMM的混合,Conformer和基于wav 2 vec 2.0的端到端电话识别器的性能和混淆矩阵。最后,未加权和加权PER与广泛的语音类先验的关系进行了研究的混合和端到端系统。
In this article, we analyse the performance of the TIMIT-based phone recognition systems beyond the overall phone error rate (PER) metric. We consider three broad phonetic classes (BPCs): {affricate, diphthong, fricative, nasal, plosive, semi-vowel, vowel, silence}, {consonant, vowel, silence} and {voiced, unvoiced, silence} and, calculate the contribution of each phonetic class in terms of the substitution, deletion, insertion and PER. Furthermore, for each BPC we investigate the following: evolution of PER during training, effect of noise (NTIMIT), importance of different spectral subbands (1, 2, 4, and 8 kHz), usefulness of bidirectional vs unidirectional sequential modelling, transfer learning from WSJ and regularisation via monophones. In addition, we construct a confusion matrix for each BPC and analyse the confusions via dimensionality reduction to 2D at the input (acoustic features) and output (logits) levels of the acoustic model. We also compare the performance and confusion matrices of the BLSTM-based hybrid baseline system with those of the GMM-HMM based hybrid, Conformer and wav2vec 2.0 based end-to-end phone recognisers. Finally, the relationship of the unweighted and weighted PERs with the broad phonetic class priors is studied for both the hybrid and end-to-end systems.