Lightweight Multi-objective Voice Adaptation for Real-time Speech Interaction Applied in Games
Lightweight Multi-objective Voice Adaptation for Real-time Speech Interaction Applied in Games
复制标题
轻量级多目标语音自适应在游戏中的实时语音交互应用
DOI:
10.1109/cog47356.2020.9231643
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Yuji Sato
中科院分区:
文献类型:
--
作者:
Mads Midtlyng; Yuji Sato
This paper proposes a novel voice adaptation method that we applied to interactive activities such as games where source and target data are unaligned. Conventional methods have seen the use of probabilistic models or more recently, Deep Neural Networks. Common for most methods is that they require multiple subjects to train in conjunction, thus voice adaptation is not practical to be used in commercial applications. We propose a method which convert audible frequencies to light spectrum simple RGB color format, and not comparing sound signal similarities, but rather likeness in color. The comparison is done using multi-objective optimization which considers raw and normalized frame colors as two separate objectives to be evaluated, respectively audible and spectral structure. The distance for the objectives is used to select an ideal output frame. Finally, prosodic information such as speech intensity is translated from measured input values onto the designated output frame. The method is evaluated using MOS, ABX, performance benchmark and lastly implemented into the Unity3D game engine as a proof of concept. Results show good sound quality and high performance with little output fragmentation.