Active listening.

Active listening.
复制标题

DOI:
10.1016/j.heares.2020.107998
复制
发表时间:
2021-01
期刊:
影响因子:
2.8
通讯作者:
Holmes E
Holmes E
中科院分区:
医学1区
文献类型:
--
作者:
Friston KJ;Sajid N;Quiroga-Martinez DR;Parr T;Price CJ;Holmes E

文献摘要

参考文献

被引文献

相似文献

本文介绍了主动倾听作为语音合成和识别的统一框架。积极倾听的概念继承自积极推理,它认为感知和行动在一个普遍的要求下:最大限度地为我们的(生成的)世界模型提供证据。首先,我们描述了一个口语生成模型,该模型模拟了(i)离散的词汇、韵律和说话者属性如何产生连续的声音信号;相反,(ii)连续的声音信号是如何被识别为单词的。“主动”方面涉及(隐蔽地)分割口语句子,并从主动视觉中借鉴思想。它将语音分割作为内部动作的选择,对应于单词边界的位置。实际上,单词边界的选择最大化了单个单词如何生成的内部模型的证据。我们通过模拟语音识别并展示句子的推断内容如何依赖于先验信念和背景噪声来建立人脸有效性。最后,我们通过将神经元或生理反应(如错配负性和P300)与积极倾听下的信念更新联系起来来考虑预测效度,这在缺乏对接下来将要听到的内容的准确先验信念的情况下是最大的。描述一个合成和识别语音的生成模型。考虑语音分割(放置词边界)作为一个积极的过程。将语音分割和词汇推理作为互补。将神经错配反应(如MMN)与信念更新联系起来。
This paper introduces active listening, as a unified framework for synthesising and recognising speech. The notion of active listening inherits from active inference, which considers perception and action under one universal imperative: to maximise the evidence for our (generative) models of the world. First, we describe a generative model of spoken words that simulates (i) how discrete lexical, prosodic, and speaker attributes give rise to continuous acoustic signals; and conversely (ii) how continuous acoustic signals are recognised as words. The ‘active’ aspect involves (covertly) segmenting spoken sentences and borrows ideas from active vision. It casts speech segmentation as the selection of internal actions, corresponding to the placement of word boundaries. Practically, word boundaries are selected that maximise the evidence for an internal model of how individual words are generated. We establish face validity by simulating speech recognition and showing how the inferred content of a sentence depends on prior beliefs and background noise. Finally, we consider predictive validity by associating neuronal or physiological responses, such as the mismatch negativity and P300, with belief updating under active listening, which is greatest in the absence of accurate prior beliefs about what will be heard next. Describes a generative model for synthesising and recognising speech. Considers speech segmentation (placing word boundaries) as an active process. Treats speech segmentation and lexical inferences as complementary. Associates neural mismatch responses (e.g., MMN) with belief updating.
DOI: 10.1016/j.neuron.2012.10.038
发表时间: 2012-11-21
期刊: Neuron
影响因子: 16.2
作者:
Bastos AM;Usrey WM;Adams RA;Mangun GR;Fries P;Friston KJ
通讯作者: Friston KJ
DOI: 10.1016/j.jmp.2015.11.003
发表时间: 2017-02
影响因子: 1.8
作者:
Bogacz R
通讯作者: Bogacz R
DOI: 10.1191/0267658305sr250oa
发表时间: 2005-10-01
影响因子: 2.4
作者:
Altenberg, EP
通讯作者: Altenberg, EP
DOI: 10.1016/j.specom.2005.02.016
发表时间: 2005-07-01
影响因子: 3.2
作者:
Bänziger, T;Scherer, KR
通讯作者: Scherer, KR
DOI: 10.1016/j.conb.2017.08.010
发表时间: 2017-10
影响因子: 5.7
作者:
Aitchison L;Lengyel M
通讯作者: Lengyel M