Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin

Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
复制标题

DOI:
--
复制
发表时间:
2015-12
期刊:
--
影响因子:
--
通讯作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;
中科院分区:
其他
文献类型:
--
作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;

文献摘要

被引文献

相似文献

我们表明,端到端的深度学习方法可用于识别英语或汉语普通话语音——这两种截然不同的语言。由于它用神经网络取代了手工设计组件的整个流程,端到端学习使我们能够处理各种各样的语音,包括嘈杂环境、口音和不同语言。我们方法的关键是应用了高性能计算(HPC)技术,使以前需要数周的实验如今能在数天内完成。这使我们能够更快地迭代,以确定更优的架构和算法。结果,在一些情况下,当在标准数据集上进行基准测试时,我们的系统与人工转录具有竞争力。最后,通过在数据中心使用一种名为GPU批处理调度的技术,我们表明我们的系统可以低成本地部署在在线环境中,在大规模服务用户时实现低延迟。
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech-two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.