Hierarchical video summarization based on context clustering

Hierarchical video summarization based on context clustering
复制标题

基于上下文聚类的层次化视频摘要

DOI:
--
复制
发表时间:
2003
期刊:
SPIE ITCom
影响因子:
--
通讯作者:
John R. Smith
John R. Smith
中科院分区:
--
文献类型:
--
作者:
Belle L. Tseng;John R. Smith

文献摘要

被引文献

相似文献

在我们的视频个性化和摘要系统中,根据用户的偏好和使用环境动态地生成个性化的视频摘要。三层个性化系统采用服务器-中间件-客户端的体系结构,以维护、选择、适配和向用户提供富媒体内容。服务器将内容源与其对应的MPEG-7元数据描述一起存储。在本文中,元数据包括视觉语义标注和自动语音转录。我们在中间件中的个性化和摘要引擎通过将镜头注释和句子转录与用户偏好相匹配来选择所需视频片段的最佳集合。除了找到想要的内容外,目标是提出一个连贯的总结。创建摘要的方法有多种,我们将重点介绍基于上下文信息生成分层视频摘要的挑战。在我们的摘要算法中,三个输入被用来生成分层的视频摘要输出。这些输入是(1)服务器中的内容的MPEG7元数据描述,(2)来自用户客户端的用户偏好和使用环境声明,以及(3)包括MPEG7控制术语列表和分类方案的上下文信息。在视频序列中,描述和相关性分数被分配给每个镜头。基于这些镜头描述,执行上下文聚类以收集连续的相似镜头以对应于分层场景表示。上下文聚类基于可用的上下文信息,并且可以从领域知识或规则引擎中得到。最后,选择结构化视频片段来生成分层摘要,有效地平衡了场景表示和镜头选择之间的关系。
A personalized video summary is dynamically generated in our video personalization and summarization system based on user preference and usage environment. The three-tier personalization system adopts the server-middleware-client architecture in order to maintain, select, adapt, and deliver rich media content to the user. The server stores the content sources along with their corresponding MPEG-7 metadata descriptions. In this paper, the metadata includes visual semantic annotations and automatic speech transcriptions. Our personalization and summarization engine in the middleware selects the optimal set of desired video segments by matching shot annotations and sentence transcripts with user preferences. Besides finding the desired contents, the objective is to present a coherent summary. There are diverse methods for creating summaries, and we focus on the challenges of generating a hierarchical video summary based on context information. In our summarization algorithm, three inputs are used to generate the hierarchical video summary output. These inputs are (1) MPEG-7 metadata descriptions of the contents in the server, (2) user preference and usage environment declarations from the user client, and (3) context information including MPEG-7 controlled term list and classification scheme. In a video sequence, descriptions and relevance scores are assigned to each shot. Based on these shot descriptions, context clustering is performed to collect consecutively similar shots to correspond to hierarchical scene representations. The context clustering is based on the available context information, and may be derived from domain knowledge or rules engines. Finally, the selection of structured video segments to generate the hierarchical summary efficiently balances between scene representation and shot selection.