Structuring lecture videos for distance learning applications

Structuring lecture videos for distance learning applications
复制标题

为远程学习应用构建讲座视频

DOI:
10.1109/mmse.2003.1254444
复制
发表时间:
2003
期刊:
Fifth International Symposium on Multimedia Software Engineering, 2003. Proceedings.
影响因子:
--
通讯作者:
T. Pong
T. Pong
中科院分区:
--
文献类型:
--
作者:
C. Ngo;Feng Wang;T. Pong

文献摘要

被引文献

相似文献

我们提出了一种自动和新颖的方法,在结构和索引远程学习应用的讲座视频。通过对视频内容进行结构化,我们可以同时支持多媒体文档的主题索引和语义查询。我们的目标是将从电子幻灯片中提取的讨论主题与其相关的视频和音频片段相关联。我们提出的方法中的两个主要技术包括视频文本分析和语音识别。最初,基于幻灯片过渡将视频划分为镜头。对于每个镜头,嵌入的视频文本被检测,重建和分割为高分辨率的前景文本,用于商业OCR识别。然后,识别的文本可以与其相关联的幻灯片进行匹配,以进行视频索引。同时,还从电子幻灯片中提取短语(标题)和关键字(内容)以识别语音信号。所发现的短语和关键字被进一步用作查询以检索最相似的幻灯片用于语音索引。
We present an automatic and novel approach in structuring and indexing lecture videos for distance learning applications. By structuring video content, we can support both topic indexing and semantic querying of multimedia documents. our aim is to link the discussion topics extracted from the electronic slides with their associated video and audio segments. Two major techniques in our proposed approach include video text analysis and speech recognition. Initially, a video is partitioned into shots based on slide transitions. For each shot, the embedded video texts are detected, reconstructed and segmented as high-resolution foreground texts for commercial OCR recognition. The recognized texts can then be matched with their associated slides for video indexing. Meanwhile, both phrases (title) and keywords (content) are also extracted from the electronic slides to spot the speech signals. The spotted phrases and keywords are further utilized as queries to retrieve the most similar slide for speech indexing.