Speech Recognition for the iCub Platform.

Speech Recognition for the iCub Platform.
复制标题

DOI:
10.3389/frobt.2018.00010
复制
发表时间:
2018
影响因子:
3.4
通讯作者:
Badino L
Badino L
中科院分区:
其他
文献类型:
--
作者:
Higy B;Mereta A;Metta G;Badino L

文献摘要

被引文献

相似文献

本文描述了用于构建自动语音识别(ASR)系统并在YARP平台上运行它们的开源软件(可在上获得)。该工具包旨在(I)允许非ASR专家轻松创建他们自己的ASR系统并在iCub上运行,以及(Ii)构建基于深度学习的模型,专门解决ASR系统在人类与iCub口头交互的背景下面临的主要挑战。该工具包主要由集成在YARP中的Python、C++代码和外壳脚本组成。作为额外的贡献,为更多想要试验生物启发和发展学习启发的ASR系统的专家ASR用户提供了第二个代码库(用MatLab编写)。具体地说,我们提供了两种不同类型的语音识别代码:“发音的”和“无监督的”语音识别。第一种理论很大程度上是受到言语知觉神经生物学理论的启发,这些理论认为言语知觉是通过大脑运动皮质活动来调节的。我们的发音系统已被证明优于强大的基于深度学习的基线。第二种类型的识别系统,“无监督”系统,不使用任何监督信息(与大多数ASR系统相反,包括我们的发音系统)。在某种程度上,他们模仿了一个必须自己发现语言的基本语音单位的婴儿。此外,我们提供了由ASR的预训练深度学习模型和2.5小时语音命令语音数据集VoCub数据集组成的资源,该数据集可用于使ASR系统适应iCub操作的典型声学环境。
This paper describes open source software (available at ) to build automatic speech recognition (ASR) systems and run them within the YARP platform. The toolkit is designed (i) to allow non-ASR experts to easily create their own ASR system and run it on iCub and (ii) to build deep learning-based models specifically addressing the main challenges an ASR system faces in the context of verbal human–iCub interactions. The toolkit mostly consists of Python, C++ code and shell scripts integrated in YARP. As additional contribution, a second codebase (written in Matlab) is provided for more expert ASR users who want to experiment with bio-inspired and developmental learning-inspired ASR systems. Specifically, we provide code for two distinct kinds of speech recognition: “articulatory” and “unsupervised” speech recognition. The first is largely inspired by influential neurobiological theories of speech perception which assume speech perception to be mediated by brain motor cortex activities. Our articulatory systems have been shown to outperform strong deep learning-based baselines. The second type of recognition systems, the “unsupervised” systems, do not use any supervised information (contrary to most ASR systems, including our articulatory systems). To some extent, they mimic an infant who has to discover the basic speech units of a language by herself. In addition, we provide resources consisting of pre-trained deep learning models for ASR, and a 2.5-h speech dataset of spoken commands, the VoCub dataset, which can be used to adapt an ASR system to the typical acoustic environments in which iCub operates.