Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions

Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions
复制标题

说话者方言、性别的影响

DOI:
--
复制
发表时间:
2017
期刊:
Interspeech
影响因子:
--
通讯作者:
C. Kasten
C. Kasten
中科院分区:
--
文献类型:
--
作者:
Rachael Tatman;C. Kasten

文献摘要

被引文献

相似文献

该项目比较了两个自动语音识别(ASR)系统的准确性-Bing Speech和YouTube的自动字幕-跨性别,种族和四种美式英语方言。所包括的方言是根据它们的声学差异而选择的。Bing语音的词汇错误率(WER)在方言和民族之间存在差异,但它们在统计上并不可靠。然而,YouTube的自动字幕在方言和种族之间确实有统计学上的差异。平均错误率最低的是普通美国人和白色谈话者。这两个系统都没有可靠的不同性别的WER,这是以前报道的YouTube的自动字幕。然而,较高的错误率非白色的谈话者是令人担忧的,因为它可能会降低这些系统的实用性的谈话者的颜色。
This project compares the accuracy of two automatic speech recognition (ASR) systems–Bing Speech and YouTube’s automatic captions–across gender, race and four dialects of American English. The dialects included were chosen for their acoustic dissimilarity. Bing Speech had differences in word error rate (WER) between dialects and ethnicities, but they were not statistically reliable. YouTube’s automatic captions, however, did have statistically different WERs between dialects and races. The lowest average error rates were for General American and white talkers, respectively. Neither system had a reliably different WER between genders, which had been previously reported for YouTube’s automatic captions [1]. However, the higher error rate non-white talkers is worrying, as it may reduce the utility of these systems for talkers of color.