Deep Learning Based Complex Spectral Mapping for Multi-Channel Speaker Separation and Speech Enhancement
Deep Learning Based Complex Spectral Mapping for Multi-Channel Speaker Separation and Speech Enhancement
批准号:
2125074
负责人:
Eric Fosler-Lussier
金额:
$39.06万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31
中文摘要
尽管基于深度学习的语音分离和自动语音识别取得了巨大的进步,但一个主要的挑战仍然是如何在存在室内混响和背景噪声的情况下分离并发说话者并识别他们的语音。本项目将开发一种多通道复杂频谱映射方法,用于多说话者分离和语音增强,以提高这种条件下的语音识别性能。该方法训练深度神经网络从复杂域的多通道输入中预测个体谈话者的实部和虚部。在将重叠的说话人分成同步流后,对说话人进行顺序分组,将同一说话人的语音在间隔内与其他说话人的语音和停顿进行分组。该方法将空间特征和频谱特征结合起来,通过多声道定位和单声道嵌入提取。递归神经网络将被训练来进行分类,以达到演讲者分类的目的,这可以处理会议中任意数量的演讲者。所提出的分离系统将使用开放的多声道扬声器分离数据集进行评估,这些数据集包含房间混响和背景噪声。该项目的研究结果有望大大提高在不利声学环境下连续说话人分离和说话人拨号的性能,有助于缩小识别单话音和识别多话音之间的性能差距。该项目的总体目标是开发一个深度学习系统,该系统可以在会话或会议设置中连续分离单个说话者,并准确识别这些说话者的话语。基于同步分组的最新进展,以独立于说话者的方式分离和增强重叠的说话者,该项目主要侧重于说话者分组,旨在将同一说话者在不同时间的话语分组。为了实现说话人的特征化,将进行基于深度学习的顺序分组,该分组将整合说话人的空间和频谱特征。通过顺序组织,同步流将与早先分离的说话人流分组形成顺序流,每个顺序流对应同一说话人截至当前时间的所有话语。将研究说话人的定位和分类,使顺序分组能够创建新的顺序流,并在会议场景中处理任意数量的说话人。通过增加空间维度,提出的拨号方法为“谁在何时何地发言”的问题提供了解决方案,大大扩展了“谁在何时发言”的传统范围。提议的分离系统将使用多通道扬声器分离数据集进行评估,这些数据集包含记录对话中的高度重叠语音,以及真实环境中存在的房间混响和背景噪声。自动语音识别的主要评价指标是单词错误率。用拨码误差率来衡量扬声器拨码的性能。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Despite tremendous advances in deep learning based speech separation and automatic speech recognition, a major challenge remains how to separate concurrent speakers and recognize their speech in the presence of room reverberation and background noise. This project will develop a multi-channel complex spectral mapping approach to multi-talker speaker separation and speech enhancement in order to improve speech recognition performance in such conditions. The proposed approach trains deep neural networks to predict the real and imaginary parts of individual talkers from the multi-channel input in the complex domain. After overlapped speakers are separated into simultaneous streams, sequential grouping will be performed for speaker diarization, which is the task of grouping the speech utterances of the same talker over intervals with the utterances of other speakers and pauses. Proposed speaker diarization will integrate spatial and spectral speaker features, which will be extracted through multi-channel speaker localization and single-channel speaker embedding. Recurrent neural networks will be trained to perform classification for the purpose of speaker diarization, which can handle an arbitrary number of speakers in a meeting. The proposed separation system will be evaluated using open, multi-channel speaker separation datasets that contain both room reverberation and background noise. The results from this project are expected to substantially elevate the performance of continuous speaker separation, as well as speaker diarization, in adverse acoustic environments, helping to close the performance gap between recognizing single-talker speech and recognizing multi-talker speech.The overall goal of this project is to develop a deep learning system that can continuously separate individual speakers in a conversational or meeting setting and accurately recognize the utterances of these speakers. Building on recent advances on simultaneous grouping to separate and enhance overlapped speakers in a talker-independent fashion, the project is mainly focused on speaker diarization, which aims to group the speech utterances of the same speaker across time. To achieve speaker diarization, deep learning based sequential grouping will be performed and it will integrate spatial and spectral speaker characteristics. Through sequential organization, simultaneous streams will be grouped with earlier-separated speaker streams to form sequential streams, each of which corresponds to all the utterances of the same speaker up to the current time. Speaker localization and classification will be investigated to make sequential grouping capable of creating new sequential streams and handling an arbitrary number of speakers in a meeting scenario. With the added spatial dimension, the proposed diarization approach provides a solution to the question of who spoke when and where, significantly expanding the traditional scope of who spoke when. The proposed separation system will be evaluated using multi-channel speaker separation datasets that contain highly overlapped speech in recorded conversations, as well as room reverberation and background noise present in real environments. The main evaluation metric will be word error rate in automatic speech recognition. The performance of speaker diarization will be measured using diarization error rate.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/taslp.2022.3202129
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[H. Taherian;Ke Tan;Deliang Wang]
通讯作者:
H. Taherian;Ke Tan;Deliang Wang
Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech Separation
用于多通道连续语音分离的多分辨率基于位置的训练
DOI:
--
发表时间:
2023
期刊:
Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing
影响因子:
--
作者:
[Hassan Taherian, DeLiang Wang]
通讯作者:
DeLiang Wang
RI: Small: Early Elementary Reading Verification in Challenging Acoustic Environments
-
批准号:2008043
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2020
-
负责人:Eric Fosler-Lussier
-
依托单位:
RI: Medium: Deep Neural Networks for Robust Speech Recognition through Integrated Acoustic Modeling and Separation
-
批准号:1409431
-
项目类别:Continuing Grant
-
资助金额:$79.81万
-
财政年份:2014
-
负责人:Eric Fosler-Lussier
-
依托单位:
CI-ADDO-NEW: Collaborative Research: The Speech Recognition Virtual Kitchen
-
批准号:1305319
-
项目类别:Standard Grant
-
资助金额:$38.21万
-
财政年份:2013
-
负责人:Eric Fosler-Lussier
-
依托单位:
CI-P:Collaborative Research:The Speech Recognition Virtual Kitchen
-
批准号:1205424
-
项目类别:Standard Grant
-
资助金额:$4.85万
-
财政年份:2012
-
负责人:Eric Fosler-Lussier
-
依托单位:
RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
-
批准号:0905420
-
项目类别:Standard Grant
-
资助金额:$33.45万
-
财政年份:2009
-
负责人:Eric Fosler-Lussier
-
依托单位:
CAREER: Breaking the phonetic code: novel acoustic-lexical modeling techniques for robust automatic speech recognition
-
批准号:0643901
-
项目类别:Continuing Grant
-
资助金额:$50.3万
-
财政年份:2006
-
负责人:Eric Fosler-Lussier
-
依托单位:
Workshop: Student Research in Computational Linguistics, at the HLT/NAACL 2004 Conference
-
批准号:0422841
-
项目类别:Standard Grant
-
资助金额:$2.02万
-
财政年份:2004
-
负责人:Eric Fosler-Lussier
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: