Ultrasound-Based Silent Speech Interface Using Convolutional and Recurrent Neural Networks

Ultrasound-Based Silent Speech Interface Using Convolutional and Recurrent Neural Networks
复制标题

使用卷积和循环神经网络的基于超声的无声语音接口

DOI:
--
复制
发表时间:
2019
影响因子:
--
通讯作者:
T. Csapó
T. Csapó
中科院分区:
物理4区
文献类型:
--
作者:
E. Juanpere;T. Csapó

文献摘要

被引文献

相似文献

无声语音接口(SSI)是一种以从发音运动合成语音为目标的技术。提出了一种基于深度神经网络的SSI,其使用舌的超声图像作为输入信号,声码器的频谱系数作为目标参数。几个深 学习模型,如基线前馈,以及卷积和递归神经网络的组合。还研究了使用深度卷积自动编码器的预处理步骤。根据实验结果,提出了一种基于CNN的结构, 双向LSTM层显示了最好的客观和主观结果。
Silent Speech Interface (SSI) is a technology with the goal of synthesizing speech from articulatory motion. A Deep Neural Network based SSI using ultrasound images of the tongue as input signals and spectral coefficients of a vocoder as target parameters are proposed. Several deep learning models, such as a baseline Feed-forward, and a combination of Convolutional and Recurrent Neural Networks are presented and discussed. A pre-processing step using a Deep Convolutional AutoEncoder was also studied. According to the experimental results, an architecture based on a CNN and bidirectional LSTM layers has shown the best objective and subjective results.