A Classification-Aided Framework for Non-Intrusive Speech Quality Assessment

A Classification-Aided Framework for Non-Intrusive Speech Quality Assessment
复制标题

DOI:
10.1109/waspaa.2019.8937192
复制
发表时间:
2019-10
期刊:
2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
影响因子:
--
通讯作者:
Xuan Dong;D. Williamson
Xuan Dong;D. Williamson
中科院分区:
其他
文献类型:
--
作者:
Xuan Dong;D. Williamson

文献摘要

被引文献

相似文献

语音质量感知评估(PESQ)等客观指标已成为评估语音的标准措施。这些指标可以实现高效且无成本的评估,通常通过将降级的语音信号与其基础的干净参考信号进行比较来计算评级。然而,基于参考的指标不能用于评估具有不可访问参考的现实世界信号。该项目开发了一个非侵入式框架,用于评估噪声和增强语音的感知质量。我们提出了一种话语级分类辅助非侵入式(UCAN)评估方法,该方法将质量得分分类任务与质量得分估计回归任务相结合。我们的方法使用分类质量排名作为辅助约束来协助质量分数估计,其中我们以多任务方式联合训练多层卷积神经网络。使用 TIMIT 语音语料库和多种信噪比下的几种噪声来评估该方法。结果表明,与几种最先进的方法相比,所提出的系统显着提高了质量分数估计。
Objective metrics, such as the perceptual evaluation of speech quality (PESQ) have become standard measures for evaluating speech. These metrics enable efficient and costless evaluations, where ratings are often computed by comparing a degraded speech signal to its underlying clean reference signal. Reference-based metrics, however, cannot be used to evaluate real-world signals that have inaccessible references. This project develops a nonintrusive framework for evaluating the perceptual quality of noisy and enhanced speech. We propose an utterance-level classification-aided non-intrusive (UCAN) assessment approach that combines the task of quality score classification with the regression task of quality score estimation. Our approach uses a categorical quality ranking as an auxiliary constraint to assist with quality score estimation, where we jointly train a multi-layered convolutional neural network in a multi-task manner. This approach is evaluated using the TIMIT speech corpus and several noises under a wide range of signal-to-noise ratios. The results show that the proposed system significantly improves quality score estimation as compared to several state-of-the-art approaches.