Control of Avatar's Facial Expression Using Fundamental Frequency in Multi-user Voice Chat System

Control of Avatar's Facial Expression Using Fundamental Frequency in Multi-user Voice Chat System
复制标题

多用户语音聊天系统中基频控制虚拟人物面部表情

DOI:
10.1007/11821830_50
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
K. Fujita
K. Fujita
中科院分区:
--
文献类型:
--
作者:
Toshimitsu Miyajima;K. Fujita

文献摘要

被引文献

相似文献

为了在多用户虚拟空间语音聊天系统中实现多人随意聊天,提出了一种基于用户话语基频的CG虚拟人物面部表情自动控制算法。该方法利用了反映情绪活动的声音基本频率的共同趋势,特别是喜悦的强度。本研究简化了面部表情控制问题,通过限制表情的喜悦强度,因为它似乎是最重要的表情,以方便随意聊天。使用基频的问题在于,基频随语调和情绪的变化而变化;因此,原始基本频率的使用热情地改变了化身的表达。为此,将情绪活动点(EPa)定义为归一化基频的移动平均值,以抑制语调的影响。基于面部动作编码系统(FACS),使用EPa对虚拟角色面部表情的愉悦程度进行线性控制。实验选择移动平均线的持续时间为5秒。然而,移动平均线会延迟化身的行为,尤其是在回应话语中的延迟更为严重。因此,为了补偿反应的延迟,我们使用反应话语的初始音量来定义反应情绪点(EPr)。EPr仅计算响应话语,即紧随另一个用户的话语之后的话语。实验确定EPr与EPa的比例为1:1。所提出的虚拟人物面部表情自动控制算法在已有的虚拟空间多用户语音聊天系统上实现。对10名受试者进行主观评价。每个被试被要求在单独的房间里与一个实验伙伴用这个系统聊天4分钟,并用李克特量表回答4个问题。在整个实验过程中,受试者对面部表情自动控制的印象较好。与固定面部表情、单独使用EPa和单独使用EPr条件下的自动控制相比,使用EPa和EPr条件下的面部控制在自然度、好感度、熟悉度和交互性方面表现更好。
An automatic facial expression control algorithm of CG avatar based on the fundamental frequency of the user’s utterance is proposed, in order to facilitate the multi-party casual chat in a multi-user virtual-space voice chat system. The proposed method utilizes the common tendency of the voice fundamental frequency that reflects the emotional activity, especially the strength of the delight. This study simplified the facial expression control problem by limiting the expression in the strength of the delight, because it appears the expression of the delight is the most important to facilitate the casual chat. The problem of using the fundamental frequency is that fundamental frequency varies with intonation as well as emotion; hence the use of the raw fundamental frequency changes the expression of the avatar passionately. Therefore, Emotional Point by emotional Activity (EPa) was defined as the moving-average of the normalized fundamental frequency, to suppress the influence of the intonation. The strength of the delight of the avatar facial expression was linearly controlled using EPa, based on the Facial Action Coding System (FACS). The duration of the moving average was chosen as five seconds experimentally. However, the moving average delays the avatar behavior, and the delay is more serious especially in the response utterance. Therefore, to compensate the delay of the response, the Emotional Point by Response (EPr), was defined using the initial voice volume of the response utterance. EPr was calculated for only the response utterance, which means the utterance just after another user’s utterance. The ratio of EPr to EPa was decided experimentally as one to one. The proposed automatic avatar facial expression control algorithm was implemented on the previously developed virtual-space multi-user voice chat system. The subjective evaluation was performed in ten subjects. The each subject in separate room was required to chat with an experimental partner using the system for four minutes and to answer four questions using Likert scale. Throughout the experiments, the subjects reported better impression of the automatic control of facial expression according to the utterances. The facial control using both EPa and EPr demonstrated better performance in terms of naturalness, favorability, familiarity and interactivity, compared to the fixed facial expression, the automatic control using EPa alone and the EPr alone conditions.