Deep neural networks for multi-channel speaker localization and speech separation
Deep neural networks for multi-channel speaker localization and speech separation
批准号:
1808932
负责人:
DeLiang Wang
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-12-01 至 2022-11-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
In recent years, there is a dramatic increase in the deployment of the voice-based interface for human-machine communication. Such devices typically have multiple microphones (or channels), and as they are used in homes, cars, and so on, a major technical challenge is how to reliably localize a target speaker and recognize his/her speech in everyday environments with multiple sound sources and room reverberation. The performance of traditional approaches to localization and separation degrades significantly in the presence of interfering sounds and room reverberation. This project investigates multi-channel speaker localization and speech separation from a deep learning perspective. The innovative approach in this project is to train deep neural networks to perform single-channel speech separation in order to identify the time-frequency regions dominated by the target speaker. Such regions across microphone pairs provide the basis for robust speaker localization and separation. Building on this novel perspective, the proposed research seeks to achieve robust speaker localization and speech separation. For robust speaker localization, time-frequency (T-F) masks will be generated by deep neural networks (DNN) from single-channel noisy speech signals. Across each pair of microphones, an integrated mask will be calculated from the two corresponding single-channel masks and then used to weight a generalized cross-correlation function, from which the direction of the target speaker will be estimated. An alternative method for localization will be based on mask-weighted steered responses. For robust speech separation, masking-based beamforming will be initially performed, where T-F masking and accurate speaker localization are expected to enhance beamforming results substantially. To overcome the limitation of spatial filtering in multi-source reverberant conditions, spectral (monaural) and spatial information will be integrated as DNN input features in order to separate only the target signal with speech characteristics and originating from a specific direction. The proposed approach will be evaluated using automatic speech recognition rate, as well as localization and separation accuracy, on multi-channel noisy and reverberant datasets recorded in real-world environments. This will ensure a broader impact not only in advancing speech processing technology but also in facilitating the design of next-generation hearing aids in the long run.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/taslp.2020.2986896
发表时间:
2020
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang]
通讯作者:
H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang
Location-based training for multi-channel talker-independent speaker separation
基于位置的多通道独立于说话者分离的训练
DOI:
--
发表时间:
2022
期刊:
Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing
影响因子:
--
作者:
[Taherian, H., Tan, K., Wang, D.L.]
通讯作者:
Wang, D.L.
DOI:
10.1109/icassp43922.2022.9746896
发表时间:
2021-07
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Zhong-Qiu Wang;Deliang Wang]
通讯作者:
Zhong-Qiu Wang;Deliang Wang
DOI:
10.1109/taslp.2022.3192104
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[H. Zhang;Deliang Wang]
通讯作者:
H. Zhang;Deliang Wang
Count and separate: incorporating speaker counting for continuous speaker separation
计数和分离:结合扬声器计数以实现连续的扬声器分离
DOI:
10.1109/icassp39728.2021.9414677
发表时间:
2021
期刊:
Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing
影响因子:
--
作者:
[Wang, Z.-Q., Wang, D.L.]
通讯作者:
Wang, D.L.
共 11 条
Collaborative Research: Separating Speech from Speech Noise to Improve Speech Intelligibility
-
批准号:0534707
-
项目类别:Standard Grant
-
资助金额:$14.49万
-
财政年份:2006
-
负责人:DeLiang Wang
-
依托单位:
ITR: Dynamics-based Speech Segregation
-
批准号:0081058
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2000
-
负责人:DeLiang Wang
-
依托单位:
Automated Auditory Scene Analysis Based on Oscillatory Correlation
-
批准号:9423312
-
项目类别:Continuing Grant
-
资助金额:$21.0万
-
财政年份:1995
-
负责人:DeLiang Wang
-
依托单位:
Segmentation and Recognition of Complex Temporal Patterns
-
批准号:9211419
-
项目类别:Continuing Grant
-
资助金额:$6.0万
-
财政年份:1992
-
负责人:DeLiang Wang
-
依托单位:
国内基金
海外基金
登录
查看更多内容
脐带间充质干细胞微囊联合低能量冲击波治疗神经损伤性ED的机制研究
-
批准号:82371631
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:卢慕峻
-
依托单位:
亚低温调控颅脑创伤急性期神经干细胞Mpc2/Lactate/H3K9lac通路促进神经修复的研究
-
批准号:82371379
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:冯军峰
-
依托单位:
基于再生运动神经路径优化Agrin作用促进损伤神经靶向投射的功能研究
-
批准号:82371373
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:沃雁
-
依托单位:
Neural Process模型的多样化高保真技术研究
-
批准号:62306326
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:王琦
-
依托单位:
声致离子电流促进小胶质细胞M2极化阻断再生神经瘢痕退变免疫机制
-
批准号:82371973
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:孙迪
-
依托单位:
生理/病理应激差异化调控肝再生的“蓝斑—中缝”神经环路机制
-
批准号:82371517
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:杨立群
-
依托单位:
LIPUS响应的弹性石墨烯多孔导管促进神经再生及其机制研究
-
批准号:82370933
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:陆家瑜
-
依托单位:
弓状核介导慢性疼痛引起动机下降的神经环路机制及rTMS干预研究
-
批准号:82371536
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:张松
-
依托单位:
听觉刺激特异性调控情绪的神经环路机制研究
-
批准号:82371516
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:周文杰
-
依托单位:
TAG1/APP信号通路调控的miRNA及其在神经前体细胞增殖和分化中的作用机制
-
批准号:31171313
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2011
-
负责人:马全红
-
依托单位: