Icassp 2022 Deep Noise Suppression Challenge

Icassp 2022 Deep Noise Suppression Challenge
复制标题

Icassp 2022 深度噪声抑制挑战赛

DOI:
10.1109/icassp43922.2022.9747230
复制
发表时间:
2022
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Sriram Srinivasan
Sriram Srinivasan
中科院分区:
--
文献类型:
--
作者:
Chandan K. A. Reddy;Harishchandra Dubey;Vishak Gopal;Ross Cutler;Sebastian Braun;H. Gamper;R. Aichner;Sriram Srinivasan

文献摘要

被引文献

相似文献

深度噪声抑制(DNS)挑战旨在促进噪声抑制领域的创新,以实现卓越的感知语音质量。这是第四届DNS挑战赛,前几届分别在InterSpeech 2020[1]、ICASSP 2021[2]和InterSpeech 2021[3]上举行。我们开源了数据集和测试集,以供研究人员训练他们的深度噪声抑制模型,以及一个基于ITU-T P.835的主观评估框架来对挑战条目进行评级和排序。我们提供对DNS-MOS P.835和单词准确性(WACC)API的访问,以挑战参与者,帮助进行迭代模型改进。在这次挑战中,我们引入了以下变化:(I)将移动设备场景包括在盲测试集中;(Ii)包括具有基线的个性化噪声抑制跟踪;(Iii)添加WACC作为客观度量;(Iv)包括DNSMOS P.835;(V)使训练数据集和测试集成为全频段(48 KHz)。我们使用WACC的平均值和主观分数P.835 SIG、BAK和OVRL来获得最终分数,以对DNS模型进行排名。我们认为,作为一个研究团体,要在具有挑战性的嘈杂现实世界场景中实现出色的语音质量,我们还有很长的路要走。
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios.