An End-To-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework

An End-To-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework
复制标题

DOI:
10.1109/icassp39728.2021.9414182
复制
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Zhuohuang Zhang;P. Vyas;Xuan Dong;D. Williamson
Zhuohuang Zhang;P. Vyas;Xuan Dong;D. Williamson
中科院分区:
其他
文献类型:
--
作者:
Zhuohuang Zhang;P. Vyas;Xuan Dong;D. Williamson

文献摘要

被引文献

相似文献

语音评估对于许多应用至关重要,但当前的侵入式方法无法在真实的环境中使用。已经提出了数据驱动的方法,但它们使用模拟语音材料或仅估计客观分数。在本文中,我们提出了一种新的多任务非侵入性的方法,能够同时估计主观和客观分数的真实世界的语音,以帮助促进学习。这种方法增强了我们之前的工作,估计主观平均意见分数,其中我们的方法现在以端到端的方式直接对时域信号进行操作。建议的系统进行比较,对几个国家的最先进的系统。实验结果表明,根据多个评估指标,我们的多任务和端到端框架可以带来更高的相关性能和更低的预测误差。
Speech assessment is crucial for many applications, but current intrusive methods cannot be used in real environments. Data-driven approaches have been proposed, but they use simulated speech materials or only estimate objective scores. In this paper, we propose a novel multi-task non-intrusive approach that is capable of simultaneously estimating both subjective and objective scores of real-world speech, to help facilitate learning. This approach enhances our prior work, which estimated subjective mean-opinion scores, where our approach now operates directly on the time-domain signal in an end-to-end fashion. The proposed system is compared against several state-of-the-art systems. The experimental results show that our multi-task and end-to-end framework leads to higher correlation performance and lower prediction errors, according to multiple evaluation measures.