Data-driven machine learning models for decoding speech categorization from evoked brain responses.

Data-driven machine learning models for decoding speech categorization from evoked brain responses.
复制标题

数据驱动的机器学习模型,用于从诱发的大脑反应中解码语音分类。

DOI:
10.1088/1741-2552/abecf0
复制
发表时间:
2021-03-23
影响因子:
4
通讯作者:
Bidelman GM
Bidelman GM
中科院分区:
工程技术2区
文献类型:
--
作者:
Mahmud MS;Yeasin M;Bidelman GM

文献摘要

参考文献

相似文献

音频的分类感知(CP)对于理解人类大脑如何感知语音声音至关重要,尽管声学特性存在广泛的变化。在这里,我们研究了听觉神经活动的时空特征,反映了CP的语音(即区分语音原型模糊的语音声音)。我们记录了64通道的脑电图作为听众迅速分类元音的声音沿着声学语音连续。我们使用支持向量机分类器和稳定性选择,以确定何时何地在大脑中的CP是最好的解码跨空间和时间通过源级分析的事件相关电位。我们发现,早期(120 ms)全脑数据解码语音类别(即原型与模糊标记)的准确率为95.16%(曲线下面积95.14%; F1评分95.00%)。对左半球(LH)和右半球(RH)反应的单独分析表明,LH解码比RH更准确和更早(89.03%对86.45%准确度; 140 ms对200 ms)。稳定性(特征)选择确定了68个脑区中的13个感兴趣区(ROI)[包括听觉皮层、缘上回和额下回(IFG)],这些脑区在刺激编码(0-260 ms)期间显示出分类表征。相比之下,15个ROI(包括额顶叶区,IFG,运动皮层)是必要的,以描述后期决策阶段(后300-800毫秒)的分类,但这些地区是高度相关的强度听者的分类听力(即斜率的行为识别功能)。我们的数据驱动的多变量模型表明,抽象类别出现令人惊讶的早期(~120毫秒)在语音处理的时间过程中,并由一个相对紧凑的额颞顶叶脑网络的参与占主导地位。
Categorical perception (CP) of audio is critical to understand how the human brain perceives speech sounds despite widespread variability in acoustic properties. Here, we investigated the spatiotemporal characteristics of auditory neural activity that reflects CP for speech (i.e. differentiates phonetic prototypes from ambiguous speech sounds). We recorded 64-channel electroencephalograms as listeners rapidly classified vowel sounds along an acoustic-phonetic continuum. We used support vector machine classifiers and stability selection to determine when and where in the brain CP was best decoded across space and time via source-level analysis of the event-related potentials. We found that early (120 ms) whole-brain data decoded speech categories (i.e. prototypical vs. ambiguous tokens) with 95.16% accuracy (area under the curve 95.14%; F1-score 95.00%). Separate analyses on left hemisphere (LH) and right hemisphere (RH) responses showed that LH decoding was more accurate and earlier than RH (89.03% vs. 86.45% accuracy; 140 ms vs. 200 ms). Stability (feature) selection identified 13 regions of interest (ROIs) out of 68 brain regions [including auditory cortex, supramarginal gyrus, and inferior frontal gyrus (IFG)] that showed categorical representation during stimulus encoding (0–260 ms). In contrast, 15 ROIs (including fronto-parietal regions, IFG, motor cortex) were necessary to describe later decision stages (later 300–800 ms) of categorization but these areas were highly associated with the strength of listeners’ categorical hearing (i.e. slope of behavioral identification functions). Our data-driven multivariate models demonstrate that abstract categories emerge surprisingly early (~120 ms) in the time course of speech processing and are dominated by engagement of a relatively compact fronto-temporal-parietal brain network.
DOI: 10.1016/j.neuroimage.2016.01.016
发表时间: 2016-04-01
期刊: NeuroImage
影响因子: 5.7
作者:
Alho J;Green BM;May PJC;Sams M;Tiitinen H;Rauschecker JP;Jääskeläinen IP
通讯作者: Jääskeläinen IP
DOI: 10.1016/j.neuropsychologia.2013.10.015
发表时间: 2014-01-01
期刊: NEUROPSYCHOLOGIA
影响因子: 2.6
作者:
Deschamps, Isabelle;Baum, Shari R.;Gracco, Vincent L.
通讯作者: Gracco, Vincent L.
DOI: 10.1016/j.brainres.2021.147385
发表时间: 2021-05-15
期刊: Brain research
影响因子: 2.9
作者:
Carter JA;Bidelman GM
通讯作者: Bidelman GM
DOI: 10.1038/ncomms12241
发表时间: 2016-08-02
影响因子: 16.6
作者:
Du Y;Buchsbaum BR;Grady CL;Alain C
通讯作者: Alain C
DOI: 10.3389/fnins.2020.00153
发表时间: 2020-02-27
影响因子: 4.3
作者:
Bidelman, Gavin M.;Bush, Lauren C.;Boudreaux, Alex M.
通讯作者: Boudreaux, Alex M.