Improving Language Identification for Multilingual Speakers

Improving Language Identification for Multilingual Speakers
复制标题

提高多语言使用者的语言识别能力

DOI:
10.1109/icassp40776.2020.9053057
复制
发表时间:
2020
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Arnab Ghoshal
Arnab Ghoshal
中科院分区:
--
文献类型:
--
作者:
Andrew R. Titus;J. Silovský;Nanxin Chen;Roger Hsiao;M. Young;Arnab Ghoshal

文献摘要

参考文献

被引文献

相似文献

近年来,口语识别(LID)技术已经从区分很大程度上不同的语言发展到区分高度相似的语言甚至同一语言的方言。然而,最被忽视的一个方面是对多语种使用者的语言歧视,尽管他们是许多利用LID技术的系统的主要目标受众。正如我们在这项工作中所展示的那样,LID系统对于大多数语言组合都具有很高的平均准确率,而当存在口音时,对于其他语言则表现不佳。我们通过使用粗粒度的目标的声学LID模型,并将其输出与上下文感知模型中的交互上下文信号,以适应每个用户的系统来解决这个问题。这个组合系统在所有语言组合中实现了平均97%的准确率,同时相对于我们的基线,最坏情况下的准确率提高了60%以上。
Spoken language identification (LID) technologies have improved in recent years from discriminating largely distinct languages to discriminating highly similar languages or even dialects of the same language. One aspect that has been mostly neglected, however, is discrimination of languages for multilingual speakers, despite being a primary target audience of many systems that utilize LID technologies. As we show in this work, LID systems can have a high average accuracy for most combinations of languages while greatly underperforming for others when accented speech is present. We address this by using coarser-grained targets for the acoustic LID model and integrating its outputs with interaction context signals in a context-aware model to tailor the system to each user. This combined system achieves an average 97% accuracy across all language combinations while improving worst-case accuracy by over 60% relative to our baseline.
多语言声学模型的神经语言代码
DOI: 10.21437/interspeech.2018-1241
发表时间: 2018
期刊:
影响因子: --
作者:
Markus Müller;Sebastian Stüker;Alex Waibel
通讯作者: Alex Waibel