Brain-to-text: decoding spoken phrases from phone representations in the brain.

Brain-to-text: decoding spoken phrases from phone representations in the brain.
复制标题

DOI:
10.3389/fnins.2015.00217
复制
发表时间:
2015
影响因子:
4.3
通讯作者:
Schultz T
Schultz T
中科院分区:
医学2区
文献类型:
--
作者:
Herff C;Heger D;de Pesters A;Telaar D;Brunner P;Schalk G;Schultz T

文献摘要

参考文献

被引文献

相似文献

长期以来,人们一直在猜测,人类和机器之间是否有可能基于自然语音相关的皮质活动进行交流。在过去的十年里,研究表明,从神经信号中识别语音的孤立方面是可行的,例如听觉特征、电话或几个孤立单词中的一个。然而,到目前为止,从与语音和语言处理相关的神经基质中解码连续的语音仍然是一个悬而未决的挑战。在这里,我们首次展示了连续说话的语音可以从脑皮层脑电(ECoG)记录中解码成表达的单词。具体地说,我们实现了一个系统,我们称之为脑到文本,模拟单个电话,使用自动语音识别(ASR)的技术,从而将说话时的大脑活动转换为相应的文本表示。实验结果表明,我们的系统可以达到25%的文字错误率和50%以下的手机错误率。此外,我们的方法通过识别那些持有关于单个电话的大量信息的皮质区域,有助于目前对连续语音产生的神经基础的理解。总之,本文描述的脑-文本系统代表着朝着基于想象语音的人机交流迈出了重要的一步。
It has long been speculated whether communication between humans and machines based on natural speech related cortical activity is possible. Over the past decade, studies have suggested that it is feasible to recognize isolated aspects of speech from neural signals, such as auditory features, phones or one of a few isolated words. However, until now it remained an unsolved challenge to decode continuously spoken speech from the neural substrate associated with speech and language processing. Here, we show for the first time that continuously spoken speech can be decoded into the expressed words from intracranial electrocorticographic (ECoG) recordings.Specifically, we implemented a system, which we call Brain-To-Text that models single phones, employs techniques from automatic speech recognition (ASR), and thereby transforms brain activity while speaking into the corresponding textual representation. Our results demonstrate that our system can achieve word error rates as low as 25% and phone error rates below 50%. Additionally, our approach contributes to the current understanding of the neural basis of continuous speech production by identifying those cortical regions that hold substantial information about individual phones. In conclusion, the Brain-To-Text system described in this paper represents an important step toward human-machine communication based on imagined speech.
DOI: 10.1371/journal.pone.0053398
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Kubanek J;Brunner P;Gunduz A;Poeppel D;Schalk G
通讯作者: Schalk G
DOI: 10.3389/fnhum.2012.00099
发表时间: 2012
影响因子: 2.9
作者:
Leuthardt EC;Pei XM;Breshears J;Gaona C;Sharma M;Freudenberg Z;Barbour D;Schalk G
通讯作者: Schalk G
DOI: 10.1007/s12021-014-9252-3
发表时间: 2015-04
期刊: Neuroinformatics
影响因子: 3
作者:
Kubanek J;Schalk G
通讯作者: Schalk G
DOI: 10.1016/j.neuroimage.2009.10.047
发表时间: 2010-02-01
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Fukuda, Miho;Rothermel, Robert;Juhasz, Csaba;Nishida, Masaaki;Sood, Sandeep;Asano, Eishi
通讯作者: Asano, Eishi
DOI: 10.1088/1741-2560/7/5/056007
发表时间: 2010-10
影响因子: 4
作者:
Kellis S;Miller K;Thomson K;Brown R;House P;Greger B
通讯作者: Greger B