An Improved Logistic Function for Mapping Raw Scores of Perceptual Evaluation of Speech Quality (PESQ)

An Improved Logistic Function for Mapping Raw Scores of Perceptual Evaluation of Speech Quality (PESQ)
复制标题

用于映射语音质量感知评估 (PESQ) 原始分数的改进逻辑函数

DOI:
--
复制
发表时间:
2018
期刊:
Journal of Engineering Research and Reports
影响因子:
--
通讯作者:
David Armando Contreras
David Armando Contreras
中科院分区:
--
文献类型:
--
作者:
A. Olatubosun;P. Olabisi;David Armando Contreras

文献摘要

被引文献

相似文献

语音服务作为电信网络的主要服务,其服务质量(QoS)水平在很大程度上决定了这些网络的性能。这项工作评估了最先进的语音质量感知评估(PESQ)客观模型,用于感知估计传输语音信号的质量。语音质量的感知估计主要由主观技术完成,结果显示为平均意见分数(MOS),其范围从1表示质量差到5表示质量好。尽管感知语音质量估计的主观方法存在局限性,但其分数可以作为将客观语音质量估计技术的质量分数相关联的基础。使用专业演播室设备和软件录制原创或参考发言,并按照ITU-T P.830的规定进行指导。演讲通过三个移动无线网络传输。建立了由64个原话(32男32女)和192个传话组成的语音数据库。参考演讲及其相应的传输(网络退化)演讲在PESQ模型上进行测试,以估计其质量分数。原始PESQ质量分数在-0.5和4.5的范围内。将其与MOS量表进行线性比较。PESQ模型的研究存在一些不足,但已有研究者对其进行了改进。评估PESQ映射函数(在ITU-T Rec P.862.1中)表明需要更好地覆盖MOS标度。对logistic增长函数的解进行了分析,并对参数进行了优化,从而开发了一种新的鲁棒logistic映射函数。原始PESQ质量分数使用开发的映射函数以及两个已知的标准映射函数进行映射,即:ITU-T P.862.1和Morfitt and Cotanis映射函数。通过三个功能得到的PESQ mos -听力质量目标(PESQ MOS-LQO)的映射分数使用方差分析进行检验,显著数字为。开发的逻辑映射功能提供了98.6%的MOS量表的质量评分覆盖率。这与两个已知的标准映射函数进行了评估,开发的函数分别比其86.8和93.7%的MOS标度覆盖率提高了11.8和4.9%。在显著性水平上,f值为60.6042,临界f值为3.04,p值为4.61721E-21。当p < 0.05时,原假设被拒绝,临界f值小于f统计量值,证实拒绝。因此,至少一个函数的数据分布具有不同的均值,并且属于单独的性能总体。
Voice service being the major offering of telecommunication networks, its level of Quality of Service (QoS) largely determines the performance of these networks. This work evaluated the state-of-the-art Perceptual Evaluation of Speech Quality (PESQ) objective model for perceptual estimation of the quality of transmitted speech signals. Perceptual estimation of the quality of speech is predominantly done by subjective techniques and the results presented as Mean Opinion Scores (MOS), which has a scale from 1 for poor quality to 5 for excellent quality. Despite constraints of the subjective approach to perceptual speech quality estimation, its scores serves as the basis for correlating quality scores from objective techniques for speech quality estimation. Original or reference speeches were recorded using professional studio equipment and software, and guided by provisions of ITU-T P.830. The speeches were transmitted over three mobile wireless networks. A speech database consisting of 64 original (32 male and 32 female) and 192 transmitted speeches was developed. Reference speeches and their corresponding transmitted (network-degraded) speeches were tested on the PESQ model to estimate their quality scores. The raw PESQ quality scores are within the scale range of -0.5 and 4.5. They were mapped to the MOS scale for linear comparison of the scales. Study of PESQ model showed several shortcomings, some of which have been improved upon by previous researchers. Evaluating PESQ mapping function (in ITU-T Rec P.862.1) showed the need for better coverage of the MOS scale. Analysis of solution for the logistic growth function was done and parameters were optimised which resulted in the development of a new robust logistic mapping function. The raw PESQ quality scores were mapped using the developed mapping function as well as two known standard mapping functions, namely: ITU-T P.862.1 and Morfitt and Cotanis mapping functions. The mapped scores known as PESQ MOS-listening quality objective (PESQ MOS-LQO) obtained with the three functions were tested using ANOVA at a significant figure of . The developed logistic mapping function offered a quality score coverage of 98.6% of the MOS scale. This was evaluated against the two known standard mapping functions and the developed function offered improvement of 11.8 and 4.9% over and above their 86.8 and 93.7% coverage of the MOS scale respectively. At the significance level of , an F-value of 60.6042, a critical-F of 3.04, and a p-value of 4.61721E-21 were obtained. With p < 0.05, the Null Hypothesis was rejected, and the critical-F value being less than the F-statistic value confirmed the rejection. Therefore, the data distribution of at least one of the functions has a different mean and belongs to a separate population of performance.