Brain-optimized extraction of complex sound features that drive continuous auditory perception

Brain-optimized extraction of complex sound features that drive continuous auditory perception
复制标题

DOI:
10.1371/journal.pcbi.1007992
复制
发表时间:
2020-07-01
影响因子:
4.3
通讯作者:
Ramsey, Nick F.
Ramsey, Nick F.
中科院分区:
生物学2区
文献类型:
--
作者:
Berezutskaya, Julia;Freudenburg, Zachary V.;Ramsey, Nick F.

文献摘要

被引文献

相似文献

了解人类大脑如何处理听觉输入仍然是一个挑战。传统上,低层次和高层次的声音特征之间的区别,但它们的定义取决于一个特定的理论框架,可能不匹配的声音的神经表征。在这里,我们假设构建一个数据驱动的听觉感知神经模型,对相关的声音特征进行最少的理论假设,可以提供一种替代方法,并可能更好地匹配神经反应。我们收集了6名患者的皮层电图记录,他们观看了一部长时间的故事片。原始电影原声被用来训练人工神经网络模型,以预测相关的神经反应。该模型实现了很高的预测准确性,并很好地推广到第二个数据集,其中新参与者观看了不同的电影。所提取的自下而上的功能捕获的声学特性是特定于声音的类型,并与各种响应延迟曲线和不同的皮层分布。具体而言,几个功能编码语音相关的声学特性,其中一些功能表现出较短的延迟曲线(与后外侧裂周皮层的响应相关),其他功能表现出较长的延迟曲线(与前外侧裂周皮层的响应相关)。我们的研究结果支持并扩展了目前的观点,通过展示存在的时间层次结构在大脑外侧裂皮层和参与的皮质网站以外的视听语音感知。
Understanding how the human brain processes auditory input remains a challenge. Traditionally, a distinction between lower- and higher-level sound features is made, but their definition depends on a specific theoretical framework and might not match the neural representation of sound. Here, we postulate that constructing a data-driven neural model of auditory perception, with a minimum of theoretical assumptions about the relevant sound features, could provide an alternative approach and possibly a better match to the neural responses. We collected electrocorticography recordings from six patients who watched a long-duration feature film. The raw movie soundtrack was used to train an artificial neural network model to predict the associated neural responses. The model achieved high prediction accuracy and generalized well to a second dataset, where new participants watched a different film. The extracted bottom-up features captured acoustic properties that were specific to the type of sound and were associated with various response latency profiles and distinct cortical distributions. Specifically, several features encoded speech-related acoustic properties with some features exhibiting shorter latency profiles (associated with responses in posterior perisylvian cortex) and others exhibiting longer latency profiles (associated with responses in anterior perisylvian cortex). Our results support and extend the current view on speech perception by demonstrating the presence of temporal hierarchies in the perisylvian cortex and involvement of cortical sites outside of this region during audiovisual speech perception.