Speech synthesis from ECoG using densely connected 3D convolutional neural networks

Speech synthesis from ECoG using densely connected 3D convolutional neural networks
复制标题

DOI:
10.1088/1741-2552/ab0c59
复制
发表时间:
2019-06-01
影响因子:
4
通讯作者:
Schultz, Tanja
Schultz, Tanja
中科院分区:
工程技术2区
文献类型:
--
作者:
Angrick, Miguel;Herff, Christian;Schultz, Tanja

文献摘要

被引文献

相似文献

Objective.从神经信号直接合成语音可以为患有神经系统疾病的人提供一种快速而自然的交流方式。无创测量的大脑活动(皮层电图; ECoG)提供了必要的时间和空间分辨率,以解码快速和复杂的过程,如语音产生。近年来,使用神经信号的语音解码已经取得了一些令人印象深刻的进展,但复杂的动力学仍然没有完全理解。然而,简单的线性模型不太可能捕捉到神经活动和连续口语之间的关系。Approach.在这里,我们展示了深度神经网络可以用于将ECoG从语音产生区域映射到语音的中间表示(logMel频谱图)。所提出的方法使用密集连接的卷积神经网络拓扑结构,非常适合处理每个参与者提供的少量数据。主要结果。在一项有六名参与者的研究中,我们在重建的和原始的logMel谱图之间实现了r = 0.69的相关性。我们通过应用Wavenet声码器将我们的预测转换回可听波形。声码器以logMel特征为条件,该特征利用更大的预先存在的数据语料库来提供最自然的声学输出。意义据我们所知,这是第一次使用深度神经网络在语音生成过程中从神经记录中重建高质量语音。
Objective. Direct synthesis of speech from neural signals could provide a fast and natural way of communication to people with neurological diseases. Invasively-measured brain activity (electrocorticography; ECoG) supplies the necessary temporal and spatial resolution to decode fast and complex processes such as speech production. A number of impressive advances in speech decoding using neural signals have been achieved in recent years, but the complex dynamics are still not fully understood. However, it is unlikely that simple linear models can capture the relation between neural activity and continuous spoken speech. Approach. Here we show that deep neural networks can be used to map ECoG from speech production areas onto an intermediate representation of speech (logMel spectrogram). The proposed method uses a densely connected convolutional neural network topology which is well-suited to work with the small amount of data available from each participant. Main results. In a study with six participants, we achieved correlations up to r = 0.69 between the reconstructed and original logMel spectrograms. We transfered our prediction back into an audible waveform by applying a Wavenet vocoder. The vocoder was conditioned on logMel features that harnessed a much larger, pre-existing data corpus to provide the most natural acoustic output. Significance. To the best of our knowledge, this is the first time that high-quality speech has been reconstructed from neural recordings during speech production using deep neural networks.