Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain

Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Divya Tadimeti;Kallirroi Georgila;D. Traum
Divya Tadimeti;Kallirroi Georgila;D. Traum
中科院分区:
其他
文献类型:
--
作者:
Divya Tadimeti;Kallirroi Georgila;D. Traum

文献摘要

被引文献

相似文献

我们评估了几个公开可用的现成的(商业和研究)自动语音识别(ASR)系统的对话代理导向的英语语音与一般美国与非美国口音的扬声器。我们的研究结果表明,非美国口音的ASR系统的性能是相当差的比一般的美国口音。根据识别器的不同,普通美国口音和所有非美国口音之间的绝对差异可以在2%到12%之间变化,相对差异在16%到49%之间变化。当我们考虑特定类别的非美国口音时,这种性能下降变得更大,这表明需要更勤奋地收集和训练非英语母语者的数据,以缩小这种性能差距。在ASR系统之间存在性能差异,虽然相同的一般模式保持不变,但对于非美国口音有更多的错误,对于某些口音,最佳识别器与整体情况不同。我们希望这些结果是有用的对话系统设计者在开发更强大的包容性对话系统,并考虑到不同口音的性能要求的ASR供应商。
We evaluate several publicly available off-the-shelf (commercial and research) automatic speech recognition (ASR) systems on dialogue agent-directed English speech from speakers with General American vs. non-American accents. Our results show that the performance of the ASR systems for non-American accents is considerably worse than for General American accents. Depending on the recognizer, the absolute difference in performance between General American accents and all non-American accents combined can vary approximately from 2% to 12%, with relative differences varying approximately between 16% and 49%. This drop in performance becomes even larger when we consider specific categories of non-American accents indicating a need for more diligent collection of and training on non-native English speaker data in order to narrow this performance gap. There are performance differences across ASR systems, and while the same general pattern holds, with more errors for non-American accents, there are some accents for which the best recognizer is different than in the overall case. We expect these results to be useful for dialogue system designers in developing more robust inclusive dialogue systems, and for ASR providers in taking into account performance requirements for different accents.