The 2020 ESPnet Update: New Features, Broadened Applications, Performance Improvements, and Future Plans

The 2020 ESPnet Update: New Features, Broadened Applications, Performance Improvements, and Future Plans
复制标题

DOI:
10.1109/dslw51110.2021.9523402
复制
发表时间:
2020-12
期刊:
2021 IEEE Data Science and Learning Workshop (DSLW)
影响因子:
--
通讯作者:
Shinji Watanabe;Florian Boyer;Xuankai Chang;Pengcheng Guo;Tomoki Hayashi;Yosuke Higuchi;Takaaki Hori;Wen-Chin Huang;H. Inaguma;Naoyuki Kamo;Shigeki Karita;Chenda Li;Jing Shi;A. Subramanian;Wangyou Zhang
Shinji Watanabe;Florian Boyer;Xuankai Chang;Pengcheng Guo;Tomoki Hayashi;Yosuke Higuchi;Takaaki Hori;Wen-Chin Huang;H. Inaguma;Naoyuki Kamo;Shigeki Karita;Chenda Li;Jing Shi;A. Subramanian;Wangyou Zhang
中科院分区:
其他
文献类型:
--
作者:
Shinji Watanabe;Florian Boyer;Xuankai Chang;Pengcheng Guo;Tomoki Hayashi;Yosuke Higuchi;Takaaki Hori;Wen-Chin Huang;H. Inaguma;Naoyuki Kamo;Shigeki Karita;Chenda Li;Jing Shi;A. Subramanian;Wangyou Zhang

文献摘要

相似文献

本文介绍了端到端语音处理工具包ESPnet(https://github.com/espnet/espnet)的最新进展。该项目于2017年12月启动,主要处理基于序列到序列建模的端到端语音识别实验。该项目发展迅速,现已涵盖广泛的语音处理应用。现在ESPnet还包括文本到语音(TTS),语音对话(VC),语音翻译(ST)和语音增强(SE),支持波束成形,语音分离,去噪和去混响。由于通用的序列到序列建模属性,所有应用程序都以端到端的方式进行训练,并且可以进一步集成和联合优化。此外,ESPnet通过整合Transformer、高级数据增强和Conformer,为这些应用提供了可再现的一体化配方,在各种基准测试中具有最先进的性能。该项目旨在为社区提供最新的语音处理经验,以便学术界和各种工业规模的研究人员可以合作开发他们的技术。
This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to mainly deal with end-to-end speech recognition experiments based on sequence-to-sequence modeling. The project has grown rapidly and now covers a wide range of speech processing applications. Now ESPnet also includes text to speech (TTS), voice conversation (VC), speech translation (ST), and speech enhancement (SE) with support for beamforming, speech separation, denoising, and dereverberation. All applications are trained in an end-to-end manner, thanks to the generic sequence to sequence modeling properties, and they can be further integrated and jointly optimized. Also, ESPnet provides reproducible all-in-one recipes for these applications with state-of-the-art performance in various benchmarks by incorporating transformer, advanced data augmentation, and conformer. This project aims to provide up-to-date speech processing experience to the community so that researchers in academia and various industry scales can develop their technologies collaboratively.