Integration of speech and vision using mutual information

Integration of speech and vision using mutual information
复制标题

使用互信息整合语音和视觉

DOI:
--
复制
发表时间:
2000
期刊:
2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
影响因子:
--
通讯作者:
D. Roy
D. Roy
中科院分区:
--
文献类型:
--
作者:
D. Roy

文献摘要

被引文献

相似文献

我们正在开发一个系统,可以从同时发生的口头和视觉输入中学习单词。目标是在没有词典的情况下在单词边界自动分割连续语音,并形成与口语单词相对应的视觉类别。互信息用于集成声学和视觉距离度量,以便从原始输入中提取视听词典。我们报告了针对婴儿的语音和图像语料库的实验结果。
We are developing a system which learns words from co-occurring spoken and visual input. The goal is to automatically segment continuous speech at word boundaries without a lexicon, and to form visual categories which correspond to spoken words. Mutual information is used to integrate acoustic and visual distance metrics in order to extract an audio-visual lexicon from raw input. We report results of experiments with a corpus of infant-directed speech and images.