Comparison of Speech Recognition Performance Between Kaldi and Google Cloud Speech API

Comparison of Speech Recognition Performance Between Kaldi and Google Cloud Speech API
复制标题

DOI:
10.1007/978-3-030-03748-2_13
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Takashi Kimura;Takashi Nose;Shinji Hirooka;Yuya Chiba;Akinori Ito
Takashi Kimura;Takashi Nose;Shinji Hirooka;Yuya Chiba;Akinori Ito
中科院分区:
其他
文献类型:
--
作者:
Takashi Kimura;Takashi Nose;Shinji Hirooka;Yuya Chiba;Akinori Ito

文献摘要

被引文献

相似文献

In recent years, many systems having a speech interface have grown. The speech interface includes spoken dialogue function and high performance of a spoken dialogue system has been required. The spoken dialogue system consists of a speech recognition module. In this study, we focus on the speech recognition module of the spoken dialogue system and aim for improving the spoken dialogue system by enhancing the performance of the speech recognition system. Among several speech recognition systems, Kaldi is a widely used speech recognition system in many kinds of researches. On the other hand, several speech recognition services that are Web API is also provided, such as IBM Watson Speech to Text, Microsoft Bing Speech API, and Google Cloud Speech API, which is known that it has high performance. This paper compares speech recognition performance between Kaldi and Google Cloud Speech API in WER and RTF and confirms the recognition performance of each recognition system.