Microscopic and Blind Prediction of Speech Intelligibility: Theory and Practice.

Microscopic and Blind Prediction of Speech Intelligibility: Theory and Practice.
复制标题

DOI:
10.1109/taslp.2022.3184888
复制
发表时间:
2022
影响因子:
5.4
通讯作者:
Kolossa, Dorothea
Kolossa, Dorothea
中科院分区:
计算机科学2区
文献类型:
--
作者:
Karbasi, Mahdie;Zeiler, Steffen;Kolossa, Dorothea

文献摘要

参考文献

被引文献

相似文献

能够在不需要听力测试的情况下估计语音可理解性,将为广泛的语音处理应用带来巨大的好处。因此,为实现这一目的,已进行了许多尝试,以引入一种客观的、理想的、没有参考资料的措施。大多数研究都是从宏观的角度来分析语音可理解度预测方法的,在较长的时间跨度内进行平均。相比之下,本文提出了一个微观评价SIP方法的理论框架。在我们的框架内,导出了基于理论的统计估计精度(StAT),它在数值上量化了微观SIP固有的统计限制。一种最先进的微观SIP方法,即使用自动语音识别(ASR)直接预测听力测试结果,在此框架内进行评估。实际结果与理论吻合较好。作为最后的贡献,引入了一个全盲的判别语音可理解性预测器(DISP),并在StAT框架内进行了评估。研究表明,这种新颖的盲估计器可以预测可理解性,甚至比基于非盲asr的方法更准确,而且其结果再次与理论推导的性能潜力很好地一致。
Being able to estimate speech intelligibility without the need for listening tests would confer great benefits for a wide range of speech processing applications. Many attempts have therefore been made to introduce an objective, and ideally referencefree measure for this purpose. Most works analyze speech intelligibility prediction (SIP) methods from a macroscopic point of view, averaging over longer time spans. This paper, in contrast, presents a theoretical framework for the microscopic evaluation of SIP methods. Within our framework, a Statistically estimated Accuracy based on Theory (StAT) is derived, which numerically quantifies the statistical limitations inherent in microscopic SIP. A state-of-the-art approach to microscopic SIP, namely, the use of automatic speech recognition (ASR) to directly predict listening test results, is evaluated within this framework. The practical results are in good agreement with the theory. As the final contribution, a fully blind DIscriminative Speech intelligibility Predictor (DISP) is introduced and is also evaluated within the StAT framework. It is shown that this novel, blind estimator can predict intelligibility as well as—and often even with better accuracy than—the non-blind ASR-based approach, and that its results are again in good agreement with its theoretically derived performance potential.
DOI: 10.1109/taslp.2016.2585878
发表时间: 2016-11-01
影响因子: 5.4
作者:
Jensen, Jesper;Taal, Cees H.
通讯作者: Taal, Cees H.
DOI: 10.3109/01050398209076203
发表时间: 1982-01-01
期刊: SCANDINAVIAN AUDIOLOGY
影响因子: --
作者:
HAGERMAN, B
通讯作者: HAGERMAN, B
DOI: 10.1109/taslp.2018.2847459
发表时间: 2018-10-01
影响因子: 5.4
作者:
Andersen, Asger Heidemann;de Haan, Jan Mark;Jensen, Jesper
通讯作者: Jensen, Jesper
DOI: 10.1177/2331216520914769
发表时间: 2020-03-01
期刊: TRENDS IN HEARING
影响因子: 2.7
作者:
Fontan, Lionel;Cretin-Maitenaz, Tom;Fullgrabe, Christian
通讯作者: Fullgrabe, Christian
DOI: 10.1121/1.3224721
发表时间: 2009-11-01
影响因子: 2.4
作者:
Juergens, Tim;Brand, Thomas
通讯作者: Brand, Thomas