SpeechQoE: A Novel Personalized QoE Assessment Model for Voice Services via Speech Sensing

SpeechQoE: A Novel Personalized QoE Assessment Model for Voice Services via Speech Sensing
复制标题

DOI:
10.1145/3560905.3568502
复制
发表时间:
2022-11
期刊:
Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems
影响因子:
--
通讯作者:
Chao Wang;Huadi Zhu;Ming Li
Chao Wang;Huadi Zhu;Ming Li
中科院分区:
其他
文献类型:
--
作者:
Chao Wang;Huadi Zhu;Ming Li

文献摘要

相似文献

体验质量(QOE)评估是一项长期但尚未解决的任务。现有的方法,尤其是用于会话语音服务的方法,被限制为利用以网络为中心的参数。然而,由于没有综合考虑与QOE相关的因素,他们的表现很难令人满意。此外,他们开发了一种对所有人都是统一的模型,因此无法处理QOE感知中的用户多样性。本文提出了一种个性化的QOE评估模型,即SPEECHQOE。它利用说话人的语音信号来推断个人在语音服务中的感知质量。SPEECHQOE从根本上解决了传统模式的缺陷。SPEECHQOE不是列举和包含无限的与QOE相关的因素,而是将固有地承载着对说话人QOE评估所需的丰富信息的语音信号作为输入。SPEECHQOE采用了一个高效的少机会学习框架,以使模型快速适应新用户。此外,我们还设计了一个轻量级的数据合成方案,以最大限度地减少模型自适应所需的数据收集开销。进一步实现了与传统参数模型的模块化集成,以避免从头开始的数据驱动方法造成的问题。我们的实验表明,SpeechQoE在QOE评估中的准确率达到了91.4%,远远超过了目前最先进的解决方案。作为这项工作的另一项贡献,我们建立了一个数据集,该数据集将成为对话呼叫的QOE评估的第一个注释音轨来源。
Quality of Experience (QoE) assessment is a long-lasting but yet-to-be-resolved task. Existing approaches, especially for conversational voice services, are restricted to leveraging network-centric parameters. However, their performances are hardly satisfactory due to the failure to consider comprehensive QoE-related factors. Moreover, they develop a one-for-all model that is uniform for all individuals and thus incapable of handling user diversity in QoE perception. This paper proposes a personalized QoE assessment model, namely SpeechQoE. It exploits speaker's speech signals to infer individual's perceived quality in voice services. SpeechQoE fundamentally addresses the drawback of conventional models. Instead of enumerating and incorporating unlimited QoE-related factors, SpeechQoE takes as input speech signals that inherently bear rich information needed for QoE assessment of the speaker. SpeechQoE employs an efficient few-shot learning framework to adapt the model to a new user quickly. We additionally design a lightweight data synthetic scheme to minimize the overhead of data collection needed for model adaption. A modular integration with a conventional parametric model is further implemented to avoid issues caused by the clean-slate data-driven approach. Our experiments show that SpeechQoE achieves an accuracy of 91.4% in QoE assessment which outperforms the state-of-the-art solutions by a clear margin. As another contribution of this work, we build a dataset that would be the first source of annotated audio tracks for QoE assessment of conversational calls.