Lightweight Multi-objective Voice Adaptation for Real-time Speech Interaction Applied in Games

Lightweight Multi-objective Voice Adaptation for Real-time Speech Interaction Applied in Games
复制标题

轻量级多目标语音自适应在游戏中的实时语音交互应用

DOI:
10.1109/cog47356.2020.9231643
复制
发表时间:
2020
期刊:
Proceedings of the 2020 IEEE Congress on Games (CoG 2020), IEEE
影响因子:
--
通讯作者:
Yuji Sato
Yuji Sato
中科院分区:
--
文献类型:
--
作者:
Mads Midtlyng; Yuji Sato

文献摘要

相似文献

本文提出了一种新的语音自适应方法,我们应用到互动活动,如游戏的源和目标数据是不一致的。传统的方法已经使用了概率模型或最近的深度神经网络。大多数方法的共同点是它们需要多个受试者一起训练,因此语音适应在商业应用中是不切实际的。我们提出了一种方法,将可听频率转换为光谱简单的RGB颜色格式,而不是比较声音信号的相似性,而是颜色的相似性。使用多目标优化来进行比较,该多目标优化将原始帧颜色和归一化帧颜色视为要评估的两个单独的目标,分别是可听结构和频谱结构。物镜的距离用于选择理想的输出帧。最后,韵律信息,如语音强度从测量的输入值转换到指定的输出帧。该方法使用MOS,ABX,性能基准进行评估,并最终实现到Unity3D游戏引擎作为概念验证。结果表明,良好的音质和高性能,输出碎片少。
This paper proposes a novel voice adaptation method that we applied to interactive activities such as games where source and target data are unaligned. Conventional methods have seen the use of probabilistic models or more recently, Deep Neural Networks. Common for most methods is that they require multiple subjects to train in conjunction, thus voice adaptation is not practical to be used in commercial applications. We propose a method which convert audible frequencies to light spectrum simple RGB color format, and not comparing sound signal similarities, but rather likeness in color. The comparison is done using multi-objective optimization which considers raw and normalized frame colors as two separate objectives to be evaluated, respectively audible and spectral structure. The distance for the objectives is used to select an ideal output frame. Finally, prosodic information such as speech intensity is translated from measured input values onto the designated output frame. The method is evaluated using MOS, ABX, performance benchmark and lastly implemented into the Unity3D game engine as a proof of concept. Results show good sound quality and high performance with little output fragmentation.