Generating Natural, Intelligible Speech From Brain Activity in Motor, Premotor, and Inferior Frontal Cortices

Generating Natural, Intelligible Speech From Brain Activity in Motor, Premotor, and Inferior Frontal Cortices
复制标题

DOI:
10.3389/fnins.2019.01267
复制
发表时间:
2019-11-22
影响因子:
4.3
通讯作者:
Schultz, Tanja
Schultz, Tanja
中科院分区:
医学2区
文献类型:
--
作者:
Herff, Christian;Diener, Lorenz;Schultz, Tanja

文献摘要

被引文献

相似文献

直接从大脑活动中产生可理解语言的神经接口将使患有严重神经障碍的人能够更自然地交流。在这里,我们记录的运动,运动前区和下额叶皮质的神经群体活动在语音生产过程中使用皮层电图(ECoG),并表明,ECoG信号单独可以用来生成可理解的语音输出,可以保存会话线索。为了直接从神经数据产生语音,我们采用了一种来自语音合成领域的方法,称为单元选择,其中语音单元被连接以形成可听输出。在我们的方法中,我们称之为Brain-to-Speech,我们根据测量的ECoG活动选择后续的语音单元,直接从神经记录中生成音频波形。Brain-To-Speech使用用户自己的声音来生成听起来非常自然的语音,并包括韵律和重音等特征。通过对参与言语产生的脑区的单独研究,我们发现言语运动皮层比其他皮层区为重建过程提供了更多的信息。
Neural interfaces that directly produce intelligible speech from brain activity would allow people with severe impairment from neurological disorders to communicate more naturally. Here, we record neural population activity in motor, premotor and inferior frontal cortices during speech production using electrocorticography (ECoG) and show that ECoG signals alone can be used to generate intelligible speech output that can preserve conversational cues. To produce speech directly from neural data, we adapted a method from the field of speech synthesis called unit selection, in which units of speech are concatenated to form audible output. In our approach, which we call Brain-To-Speech, we chose subsequent units of speech based on the measured ECoG activity to generate audio waveforms directly from the neural recordings. Brain-To-Speech employed the user's own voice to generate speech that sounded very natural and included features such as prosody and accentuation. By investigating the brain areas involved in speech production separately, we found that speech motor cortex provided more information for the reconstruction process than the other cortical areas.