FaceSync: A Linear Operator for Measuring Synchronization of Video Facial Images and Audio Tracks

FaceSync: A Linear Operator for Measuring Synchronization of Video Facial Images and Audio Tracks
复制标题

DOI:
--
复制
发表时间:
2000
影响因子:
15.1
通讯作者:
M. Slaney;M. Covell
M. Slaney;M. Covell
中科院分区:
工程技术1区
文献类型:
--
作者:
M. Slaney;M. Covell

文献摘要

被引文献

相似文献

FaceSync是一种最佳线性算法,可以找到人类说话者的音频和图像记录之间的同步程度。使用典型相关,它找到最佳方向来联合收割机组合所有音频和图像数据,将它们投影到单个轴上。FaceSync使用Pearson相关性来衡量音频和图像数据之间的同步程度。我们推导出最佳的线性变换联合收割机结合的音频和视频信息,并描述了一种实现,避免了计算相关矩阵所造成的数值问题。
FaceSync is an optimal linear algorithm that finds the degree of synchronization between the audio and image recordings of a human speaker. Using canonical correlation, it finds the best direction to combine all the audio and image data, projecting them onto a single axis. FaceSync uses Pearson's correlation to measure the degree of synchronization between the audio and image data. We derive the optimal linear transform to combine the audio and visual information and describe an implementation that avoids the numerical problems caused by computing the correlation matrices.