The Pytorch-kaldi Speech Recognition Toolkit

The Pytorch-kaldi Speech Recognition Toolkit
复制标题

DOI:
10.1109/icassp.2019.8683713
复制
发表时间:
2018-11
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
M. Ravanelli;Titouan Parcollet;Yoshua Bengio
M. Ravanelli;Titouan Parcollet;Yoshua Bengio
中科院分区:
其他
文献类型:
--
作者:
M. Ravanelli;Titouan Parcollet;Yoshua Bengio

文献摘要

被引文献

相似文献

开源软件的可用性在语音识别和深度学习的普及中发挥着显着的作用。例如,Kaldi现在是一个用于开发最先进的语音识别器的既定框架。PyTorch用于使用Python语言构建神经网络,由于其简单性和灵活性,最近在机器学习社区中引起了极大的兴趣。PyTorch-Kaldi项目旨在弥合这些流行工具包之间的差距,试图继承Kaldi的效率和PyTorch的灵活性。PyTorch-Kaldi不仅是这些软件之间的一个简单接口,而且它嵌入了几个用于开发现代语音识别器的有用功能。例如,代码是专门设计的,可以自然地插入用户定义的声学模型。作为替代方案,用户可以利用几个预先实现的神经网络,这些神经网络可以使用直观的配置文件进行自定义。PyTorch-Kaldi支持多个特征和标签流以及神经网络的组合,从而能够使用复杂的神经架构。该工具包是公开发布的沿着丰富的文档,旨在在本地或HPC集群上正常工作。在多个数据集和任务上进行的实验表明,PyTorch-Kaldi可以有效地用于开发现代最先进的语音识别器。
The availability of open-source software is playing a remarkable role in the popularization of speech recognition and deep learning. Kaldi, for instance, is nowadays an established framework used to develop state-of-the-art speech recognizers. PyTorch is used to build neural networks with the Python language and has recently spawn tremendous interest within the machine learning community thanks to its simplicity and flexibility.The PyTorch-Kaldi project aims to bridge the gap between these popular toolkits, trying to inherit the efficiency of Kaldi and the flexibility of PyTorch. PyTorch-Kaldi is not only a simple interface between these software, but it embeds several useful features for developing modern speech recognizers. For instance, the code is specifically designed to naturally plug-in user-defined acoustic models. As an alternative, users can exploit several pre-implemented neural networks that can be customized using intuitive configuration files. PyTorch-Kaldi supports multiple feature and label streams as well as combinations of neural networks, enabling the use of complex neural architectures. The toolkit is publicly-released along with a rich documentation and is designed to properly work locally or on HPC clusters.Experiments, that are conducted on several datasets and tasks, show that PyTorch-Kaldi can effectively be used to develop modern state-of-the-art speech recognizers.