Reducing Bias in Production Speech Models

Reducing Bias in Production Speech Models
复制标题

减少生产语音模型中的偏差

DOI:
--
复制
发表时间:
2017
期刊:
arXiv.org
影响因子:
--
通讯作者:
Zhenyao Zhu
Zhenyao Zhu
中科院分区:
--
文献类型:
--
作者:
Eric Battenberg;R. Child;Adam Coates;Christopher Fougner;Yashesh Gaur;Jiaji Huang;Heewoo Jun;Ajay Kannan;Markus Kliegl;Atul Kumar;Hairong Liu;Vinay Rao;S. Satheesh;David Seetapun;Anuroop Sriram;Zhenyao Zhu

文献摘要

被引文献

相似文献

用端到端深度学习系统取代手工设计的管道,在语音和对象识别等应用中取得了强劲的成果。然而,生产系统的因果关系和延迟限制使端到端语音模型重新陷入欠拟合状态,并暴露出模型中的偏差,而我们表明这些偏差无法通过“扩展”来克服,即在更多数据上训练更大的模型。在这项工作中,我们系统地识别并解决偏差来源,将错误率降低高达 20%,同时保持部署的实用性。我们通过利用改进的神经架构进行流推理、解决优化问题以及采用提高音频和标签建模多功能性的策略来实现这一目标。
Replacing hand-engineered pipelines with end-to-end deep learning systems has enabled strong results in applications like speech and object recognition. However, the causality and latency constraints of production systems put end-to-end speech models back into the underfitting regime and expose biases in the model that we show cannot be overcome by "scaling up", i.e., training bigger models on more data. In this work we systematically identify and address sources of bias, reducing error rates by up to 20% while remaining practical for deployment. We achieve this by utilizing improved neural architectures for streaming inference, solving optimization issues, and employing strategies that increase audio and label modelling versatility.