RI: Medium: Neuromorphic and Data-Driven Speech Segregation
RI: Medium: Neuromorphic and Data-Driven Speech Segregation
批准号:
1764010
负责人:
Shihab Shamma
金额:
$85.17万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-10-01 至 2023-09-30
中文摘要
该项目研究了大脑皮层中语音和音乐的神经表征是如何适应并应用于克服在极其嘈杂和混乱的环境中强大感知的挑战,模拟大脑的处理和能力。更具体地说,该项目将制定受大脑结构启发的算法,以隔离和跟踪目标扬声器或声源,测试它们的性能,并将它们与利用深度人工神经网络完成这些任务的最先进方法联系起来。使用这些算法的人类心理声学和生理实验将进行,以测试这些模仿人类能力的想法的有效性。这一努力将促进以大脑及其认知功能为模型的新型神经形态计算工具的发展。反过来,这些将提供一个理论框架来指导未来的实验,以了解复杂的认知功能是如何产生的,以及它们如何影响感官知觉并导致稳健的行为表现。计划中的项目将分为两种形式。第一个尝试借鉴现有的依赖于皮层表征的神经形态学方法,在深度神经网络框架内开发新的嵌入,这将反过来赋予后者在挑战意外环境中的类似大脑的鲁棒性。在这方面将进行三个具体的努力:使用语音和音乐的皮质表示学习深度神经网络嵌入,使用对抗性自编码器探索皮质特征的无监督聚类,以及利用音高和音色表示来增强声音的分离。第二种风格的项目借鉴了深度神经网络方法,将通过在可用数据库上进行训练而获得的理想性能和灵活性构建到神经形态算法中。计划开展两个广泛的研究领域:一个领域侧重于从深度神经网络工具箱和思想中受益的神经形态实现问题,特别是在分离和重建方面。另一个重点是研究如何利用自动编码器来有效地实现特征缩减和聚类。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project investigates how neural representations of speech and music in the cortex can be adapted and applied to overcome the challenge of robust perception in extremely noisy and cluttered environments, mimicking processing and capabilities of the brain. More specifically, the project will formulate algorithms inspired by the architecture of the brain to segregate and track targeted speakers or sound sources, test their performance, and relate them to state-of-the-art approaches that utilize deep artificial neural networks to accomplish these tasks. Human psychoacoustic and physiological experiments with these algorithms will be conducted to test the validity of these ideas for mimicking human abilities. This effort will spur the development of new neuromorphic computational tools modeled after the brain and its cognitive functions. In turn, these will provide a theoretical framework to guide future experiments into how complex cognitive functions originate and how they influence sensory perception and lead to robust behavioral performance.The planned projects will be organized into two flavors. The first attempts to borrow from existing neuromorphic approaches that rely on cortical representations to develop new embeddings within the deep neural networks framework, which will in turn endow the latter with brain-like robustness in challenging unanticipated environments. Three specific efforts within this flavor will be conducted: Learning DNN embeddings using cortical representations of speech and music, exploring unsupervised clustering of cortical features using adversarial auto-encoders, and exploiting pitch and timbre representations to enhance segregation of sound. The second flavor of projects borrows from the DNN approach to build into neuromorphic algorithms the desirable performance and flexibility attained by training on available databases. Two broad areas of studies are planned: one focuses on questions of neuromorphic implementations that benefit from DNN toolboxes and ideas, especially in segregation and reconstruction. The other focuses on investigating how autoencoders can be exploited to implement feature reduction and clustering efficiently.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systems
和谐性在基于 DNN 的系统与受生物启发的单耳语音分离系统中发挥着关键作用
DOI:
10.1109/icassp43922.2022.9747314
发表时间:
2022
期刊:
Speech and Signal Processing
影响因子:
--
作者:
[Parikh, Rahil, Kavalerov, Ilya, Espy-Wilson, Carol, Shamma, Shihab]
通讯作者:
Shamma, Shihab
The Mirrornet : Learning Audio Synthesizer Controls Inspired by Sensorimotor Interaction
镜网:受感觉运动交互启发学习音频合成器控制
DOI:
10.1109/icassp43922.2022.9747358
发表时间:
2022
期刊:
Speech and Signal Processing
影响因子:
--
作者:
[Siriwardena, Yashish M., Marion, Guilhem, Shamma, Shihab]
通讯作者:
Shamma, Shihab
Acoustic To Articulatory Speech Inversion Using Multi-Resolution Spectro-Temporal Representations Of Speech Signals
使用语音信号的多分辨率时谱表示的声学到发音语音反演
DOI:
--
发表时间:
2022
期刊:
Interspeech
影响因子:
--
作者:
[R Parikh, N Seneviratne]
通讯作者:
R Parikh, N Seneviratne
Unsupervised speaker adaptation for speaker independent acoustic to articulatory speech inversion
用于独立于说话人的声学到发音语音反转的无监督说话人自适应
DOI:
10.1121/1.5116130
发表时间:
2019
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Sivaraman, Ganesh, Mitra, Vikramjit, Nam, Hosung, Tiede, Mark, Espy-Wilson, Carol]
通讯作者:
Espy-Wilson, Carol
Collaborative Research: The computational and neural basis of statistical learning during musical enculturation
-
批准号:2242085
-
项目类别:Standard Grant
-
资助金额:$23.2万
-
财政年份:2023
-
负责人:Shihab Shamma
-
依托单位:
The Neural and Social Bases of Creative Movement
-
批准号:2024837
-
项目类别:Continuing Grant
-
资助金额:$4.99万
-
财政年份:2020
-
负责人:Shihab Shamma
-
依托单位:
Annual Telluride Workshop on Neuromorphic Engineering
-
批准号:0097975
-
项目类别:Continuing Grant
-
资助金额:$18.0万
-
财政年份:2001
-
负责人:Shihab Shamma
-
依托单位:
Workshop: Annual Telluride Workshop on Neuromorphic Engineering: June 29 thru July 19, 1998: Telluride, CO
-
批准号:9803836
-
项目类别:Standard Grant
-
资助金额:$18.58万
-
财政年份:1998
-
负责人:Shihab Shamma
-
依托单位:
Design and Fabrication of Neural Networks for Signal Processing Recognition
-
批准号:8716099
-
项目类别:Continuing Grant
-
资助金额:$22.78万
-
财政年份:1988
-
负责人:Shihab Shamma
-
依托单位:
Research Initiation: Schemes for the Analysis and Recognit-ion of Speech Based on the Fundamental Principles of Sound Processing in the Auditory System
-
批准号:8505581
-
项目类别:Standard Grant
-
资助金额:$5.8万
-
财政年份:1985
-
负责人:Shihab Shamma
-
依托单位:
海外基金