课题基金 / 基金详情

Deep neural networks for multi-channel speaker localization and speech separation

Deep neural networks for multi-channel speaker localization and speech separation
用于多通道说话者定位和语音分离的深度神经网络
批准号:
1808932
负责人:
DeLiang Wang
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-12-01 至 2022-11-30

项目摘要

项目成果

DeLiang Wang的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In recent years, there is a dramatic increase in the deployment of the voice-based interface for human-machine communication. Such devices typically have multiple microphones (or channels), and as they are used in homes, cars, and so on, a major technical challenge is how to reliably localize a target speaker and recognize his/her speech in everyday environments with multiple sound sources and room reverberation. The performance of traditional approaches to localization and separation degrades significantly in the presence of interfering sounds and room reverberation. This project investigates multi-channel speaker localization and speech separation from a deep learning perspective. The innovative approach in this project is to train deep neural networks to perform single-channel speech separation in order to identify the time-frequency regions dominated by the target speaker. Such regions across microphone pairs provide the basis for robust speaker localization and separation. Building on this novel perspective, the proposed research seeks to achieve robust speaker localization and speech separation. For robust speaker localization, time-frequency (T-F) masks will be generated by deep neural networks (DNN) from single-channel noisy speech signals. Across each pair of microphones, an integrated mask will be calculated from the two corresponding single-channel masks and then used to weight a generalized cross-correlation function, from which the direction of the target speaker will be estimated. An alternative method for localization will be based on mask-weighted steered responses. For robust speech separation, masking-based beamforming will be initially performed, where T-F masking and accurate speaker localization are expected to enhance beamforming results substantially. To overcome the limitation of spatial filtering in multi-source reverberant conditions, spectral (monaural) and spatial information will be integrated as DNN input features in order to separate only the target signal with speech characteristics and originating from a specific direction. The proposed approach will be evaluated using automatic speech recognition rate, as well as localization and separation accuracy, on multi-channel noisy and reverberant datasets recorded in real-world environments. This will ensure a broader impact not only in advancing speech processing technology but also in facilitating the design of next-generation hearing aids in the long run.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/taslp.2020.2986896
发表时间: 2020
期刊: IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子: --
作者: [H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang]
通讯作者: H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang
Location-based training for multi-channel talker-independent speaker separation
基于位置的多通道独立于说话者分离的训练
DOI: --
发表时间: 2022
期刊: Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing
影响因子: --
作者: [Taherian, H., Tan, K., Wang, D.L.]
通讯作者: Wang, D.L.
DOI: 10.1109/icassp43922.2022.9746896
发表时间: 2021-07
期刊: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者: [Zhong-Qiu Wang;Deliang Wang]
通讯作者: Zhong-Qiu Wang;Deliang Wang
DOI: 10.1109/taslp.2022.3192104
发表时间: 2022
期刊: IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子: --
作者: [H. Zhang;Deliang Wang]
通讯作者: H. Zhang;Deliang Wang
11
    Collaborative Research: Separating Speech from Speech Noise to Improve Speech Intelligibility
    ITR: Dynamics-based Speech Segregation
    Automated Auditory Scene Analysis Based on Oscillatory Correlation
    Segmentation and Recognition of Complex Temporal Patterns
    国内基金
    海外基金
    脐带间充质干细胞微囊联合低能量冲击波治疗神经损伤性ED的机制研究
    • 批准号:
      82371631
    • 项目类别:
      面上项目
    • 资助金额:
      49.00万元
    • 批准年份:
      2023
    • 负责人:
      卢慕峻
    • 依托单位:
    亚低温调控颅脑创伤急性期神经干细胞Mpc2/Lactate/H3K9lac通路促进神经修复的研究
    • 批准号:
      82371379
    • 项目类别:
      面上项目
    • 资助金额:
      49.00万元
    • 批准年份:
      2023
    • 负责人:
      冯军峰
    • 依托单位:
    基于再生运动神经路径优化Agrin作用促进损伤神经靶向投射的功能研究
    • 批准号:
      82371373
    • 项目类别:
      面上项目
    • 资助金额:
      49.00万元
    • 批准年份:
      2023
    • 负责人:
      沃雁
    • 依托单位:
    Neural Process模型的多样化高保真技术研究