Learning audio-visual associations using mutual information

Learning audio-visual associations using mutual information
复制标题

使用互信息学习视听关联

DOI:
10.1109/isiu.1999.824909
复制
发表时间:
1999
期刊:
Proceedings Integration of Speech and Image Understanding
影响因子:
--
通讯作者:
A. Pentland
A. Pentland
中科院分区:
--
文献类型:
--
作者:
D. Roy;B. Schiele;A. Pentland

文献摘要

被引文献

相似文献

本文讨论的问题,找到有用的音频和视觉输入信号之间的关联。该方法是基于最大化的互信息的视听集群。这种方法导致连续语音信号的分割,并找到对应于分割的口语单词的视觉类别。这样的视听关联可以用于对婴儿语言习得进行建模,并且动态地个性化用于包括目录浏览和可穿戴计算的各种应用的基于语音的人机界面。本文介绍了一种实现的系统,学习形状的名称,从摄像头和麦克风输入。我们目前的结果在建模语言学习领域的系统的评估。
This paper addresses the problem of finding useful associations between audio and visual input signals. The proposed approach is based on the maximization of mutual information of audio-visual clusters. This approach results in segmentation of continuous speech signals, and finds visual categories which correspond to segmented spoken words. Such audio-visual associations may be used for modeling infant language acquisition and to dynamically personalize speech-based human-computer interfaces for various applications including catalog browsing and wearable computing. This paper describes an implemented system for learning shape names from camera and microphone input. We present results in an evaluation of the system for the domain of modeling language learning.