Reinforcement Learning of Speech Recognition System Based on Policy Gradient and Hypothesis Selection

Reinforcement Learning of Speech Recognition System Based on Policy Gradient and Hypothesis Selection
复制标题

DOI:
10.1109/icassp.2018.8462656
复制
发表时间:
2017-11
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Taku Kato;T. Shinozaki
Taku Kato;T. Shinozaki
中科院分区:
其他
文献类型:
--
作者:
Taku Kato;T. Shinozaki

文献摘要

相似文献

自动语音识别(ASR)系统在许多任务上都取得了很高的识别性能。然而,这种系统的性能依赖于为监督训练准备大量任务匹配的转录语音数据的极其昂贵的开发工作。这里的关键问题是转录语音数据的成本。为了支持新的语言和新的任务,成本是不断增加的。假设为许多用户转录语音数据的广泛网络服务,如果系统具有从用户非常轻微的反馈中学习而不会惹恼他们的能力,那么它将变得更加自给自足和更有用。在本文中,我们提出了一种基于策略梯度方法的通用ASR系统强化学习框架。作为该框架的一个特殊实例,我们还提出了一种基于假设选择的强化学习方法。提出的框架为现有的几种训练和适应方法提供了新的视角。实验结果表明,与无监督自适应相比,该方法提高了识别性能。
Automatic speech recognition (ASR) systems have achieved high recognition performance for several tasks. However, the performance of such systems is dependent on the tremendously costly development work of preparing vast amounts of task-matched transcribed speech data for supervised training. The key problem here is the cost of transcribing speech data. The cost is repeatedly required to support new languages and new tasks. Assuming broad network services for transcribing speech data for many users, a system would become more self-sufficient and more useful if it possessed the ability to learn from very light feedback from the users without annoying them. In this paper, we propose a general reinforcement learning framework for ASR systems based on the policy gradient method. As a particular instance of the framework, we also propose a hypothesis selection-based reinforcement learning method. The proposed framework provides a new view for several existing training and adaptation methods. The experimental results show that the proposed method improves the recognition performance compared to unsupervised adaptation.