Automatic multimedia indexing: combining audio, speech, and visual information to index broadcast news

Automatic multimedia indexing: combining audio, speech, and visual information to index broadcast news
复制标题

自动多媒体索引:结合音频、语音和视觉信息来索引广播新闻

DOI:
10.1109/msp.2006.1621450
复制
发表时间:
2006
影响因子:
14.9
通讯作者:
Y. Hayashi
Y. Hayashi
中科院分区:
工程技术1区
文献类型:
--
作者:
K. Ohtsuki;K. Bessho;Y. Matsuo;S. Matsunaga;Y. Hayashi

文献摘要

被引文献

相似文献

本文介绍了一个索引系统,自动创建元数据的多媒体广播新闻内容,通过整合音频,语音和视觉信息。自动多媒体内容索引系统包括声学分段(AS)、自动语音识别(ASR)、主题分段(TS)和视频索引特征。AS模块中新的基于频谱的特征和平滑方法提高了从输入新闻内容的音频流中的语音检测性能。在语音识别模块中,声学模型的自动选择实现了低WER(如使用多个声学模型的并行识别)和快速识别(如使用单个声学模型)。使用词概念向量的TS方法比使用局部词频向量的传统方法获得了更准确的结果。信息集成模块提供集成来自AS模块、TS模块和SC模块的结果的功能。将其与AS结果和SC结果相结合,与单独的TS结果相比,提高了层边界检测的准确性
This paper describes an indexing system that automatically creates metadata for multimedia broadcast news content by integrating audio, speech, and visual information. The automatic multimedia content indexing system includes acoustic segmentation (AS), automatic speech recognition (ASR), topic segmentation (TS), and video indexing features. The new spectral-based features and smoothing method in the AS module improved the speech detection performance from the audio stream of the input news content. In the speech recognition module, automatic selection of acoustic models achieved both a low WER, as with parallel recognition using multiple acoustic models, and fast recognition, as with the single acoustic model. The TS method using word concept vectors achieved more accurate results than the conventional method using local word frequency vectors. The information integration module provides the functionality of integrating results from the AS module, TS module, and SC module. The story boundary detection accuracy was improved by combining it with the AS results and the SC results compared to the sole TS results