A real-time speech-driven talking head using active appearance models

A real-time speech-driven talking head using active appearance models
复制标题

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
B. Theobald;N. Wilkinson
B. Theobald;N. Wilkinson
中科院分区:
其他
文献类型:
--
作者:
B. Theobald;N. Wilkinson

文献摘要

相似文献

在本文中,我们描述了一种实时语音驱动的方法,用于合成逼真的视频序列的主题阐述任意短语。在离线训练阶段中,主动外观模型(AAM)是从手工标记的图像中构建的,并用于对背诵一些训练句子的受试者的面部进行编码。典型相关分析(CCA)结合线性回归,然后使用模型之间的关系听觉和视觉功能,这是后来用来预测视觉功能的听觉功能,为新的话语。我们目前的结果进行的实验:1)确定用于基于AAM的语音驱动的说话头中的若干听觉特征的适用性,2)确定训练集的大小对听觉和视觉特征之间的相关性的影响,3)确定上下文对相关程度的影响,以及4)确定应当从其计算听觉特征的适当窗口大小。这种方法显示出了希望,更长期的目标是开发一个完全表达的三维说话头。
In this paper we describe a real-time speech-driven method for synthesising realistic video sequences of a subject enunciating arbitrary phrases. In an offline training phase an active appearance model (AAM) is constructed from hand-labelled images and is used to encode the face of a subject reciting a few training sentences. Canonical correlation analysis (CCA) coupled with linear regression is then used to model the relationship between auditory and visual features, which is later used to predict visual features from the auditory features for novel utterances. We present results from experiments conducted: 1) to determine the suitability of several auditory features for use in an AAM-based speech-driven talking head, 2) to determine the effect of the size of the training set on the correlation between the auditory and visual features, 3) to determine the influence of context on the degree of correlation, and 4) to determine the appropriate window size from which the auditory features should be calculated. This approach shows promise and a longer term goal is to develop a fully expressive, three-dimensional talking head.