A Framework for Speech Activity Detection Using Adaptive Auditory Receptive Fields.

A Framework for Speech Activity Detection Using Adaptive Auditory Receptive Fields.
复制标题

DOI:
10.1109/taslp.2015.2481179
复制
发表时间:
2015-12
期刊:
IEEE/ACM transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Elhilali M
Elhilali M
中科院分区:
其他
文献类型:
--
作者:
Carlin MA;Elhilali M

文献摘要

被引文献

相似文献

大脑声音处理的标志之一是神经系统适应不断变化的行为需求和周围声景的能力。它可以动态地转移感官和认知资源,以专注于相关的声音。神经生理学研究表明,这种能力是通过自适应地重新调整皮质频谱-时间感受野(STRF)的形状来增强目标声音的特征,同时抑制与任务无关的干扰因素的。由于人类交流的一个重要组成部分是听者在嘈杂环境中动态跟踪语音的能力,因此听觉神经生理学获得的解决方案意味着语音活动检测(SAD)的有用适应策略。 SAD 是许多自动语音处理系统中重要的第一步,在高噪声环境中性能通常会降低。在本文中,我们描述了如何在神经生理学 STRF 集合中诱导任务驱动的适应,并展示语音适应的 STRF 如何重新定向以增强语音的频谱时间调制,同时抑制与各种非语音相关的调制。然后,我们展示了与未适应的集合和抗噪声基线相比,适应的 STRF 集合如何能够更好地检测看不见的噪声环境中的语音。最后,我们使用刺激重建任务来演示适应的 STRF 集成如何更好地捕获干净和嘈杂条件下有人参与的语音的频谱时间调制。我们的结果表明,生物学上合理的适应框架可以应用于语音处理系统,以动态地适应特征表示,从而提高噪声鲁棒性。
One of the hallmarks of sound processing in the brain is the ability of the nervous system to adapt to changing behavioral demands and surrounding soundscapes. It can dynamically shift sensory and cognitive resources to focus on relevant sounds. Neurophysiological studies indicate that this ability is supported by adaptively retuning the shapes of cortical spectro-temporal receptive fields (STRFs) to enhance features of target sounds while suppressing those of task-irrelevant distractors. Because an important component of human communication is the ability of a listener to dynamically track speech in noisy environments, the solution obtained by auditory neurophysiology implies a useful adaptation strategy for speech activity detection (SAD). SAD is an important first step in a number of automated speech processing systems, and performance is often reduced in highly noisy environments. In this paper, we describe how task-driven adaptation is induced in an ensemble of neurophysiological STRFs, and show how speech-adapted STRFs reorient themselves to enhance spectro-temporal modulations of speech while suppressing those associated with a variety of nonspeech sounds. We then show how an adapted ensemble of STRFs can better detect speech in unseen noisy environments compared to an unadapted ensemble and a noise-robust baseline. Finally, we use a stimulus reconstruction task to demonstrate how the adapted STRF ensemble better captures the spectrotemporal modulations of attended speech in clean and noisy conditions. Our results suggest that a biologically plausible adaptation framework can be applied to speech processing systems to dynamically adapt feature representations for improving noise robustness.