Audio-Visual Group Recognition Using Diffusion Maps

Audio-Visual Group Recognition Using Diffusion Maps
复制标题

使用扩散图进行视听组识别

DOI:
--
复制
发表时间:
2010
影响因子:
5.4
通讯作者:
S. Zucker
S. Zucker
中科院分区:
工程技术1区
文献类型:
--
作者:
Y. Keller;R. Coifman;Stéphane Lafon;S. Zucker

文献摘要

参考文献

被引文献

相似文献

数据融合是恢复物理系统状态的自然且常见的方法。但不同传感器的不同外观仍然是一个根本障碍。我们提出了一种基于光谱扩散框架的多感官数据的统一嵌入方案,解决了这个问题。我们的方案纯粹是数据驱动的,并且假设没有数据源的先验统计或确定性模型。为了提取底层结构,我们首先单独嵌入每个输入通道;然后将所得结构组合在扩散坐标中。特别是,当不同的传感器以不同的采样密度采样相似的现象时,我们应用密度不变的 Laplace-Beltrami 嵌入。这是多传感器采集和处理中的一个基本问题,在先前的方法中被忽视。我们扩展了之前关于群体识别的工作,并提出了一种选择扩散坐标的新方法。为了验证我们的方法,我们展示了音频/视觉语音识别的性能改进。
Data fusion is a natural and common approach to recovering the state of physical systems. But the dissimilar appearance of different sensors remains a fundamental obstacle. We propose a unified embedding scheme for multisensory data, based on the spectral diffusion framework, which addresses this issue. Our scheme is purely data-driven and assumes no a priori statistical or deterministic models of the data sources. To extract the underlying structure, we first embed separately each input channel; the resultant structures are then combined in diffusion coordinates. In particular, as different sensors sample similar phenomena with different sampling densities, we apply the density invariant Laplace-Beltrami embedding. This is a fundamental issue in multisensor acquisition and processing, overlooked in prior approaches. We extend previous work on group recognition and suggest a novel approach to the selection of diffusion coordinates. To verify our approach, we demonstrate performance improvements in audio/visual speech recognition.
DOI: 10.1073/pnas.0500334102
发表时间: 2005-05-24
影响因子: 11.1
作者:
Coifman, RR;Lafon, S;Zucker, SW
通讯作者: Zucker, SW