Visual Sentences for Pose Retrieval Over Low-Resolution Cross-Media Dance Collections

Visual Sentences for Pose Retrieval Over Low-Resolution Cross-Media Dance Collections
复制标题

DOI:
10.1109/tmm.2012.2199971
复制
发表时间:
2012-12
影响因子:
7.3
通讯作者:
Reede Ren;J. Collomosse
Reede Ren;J. Collomosse
中科院分区:
计算机科学1区
文献类型:
--
作者:
Reede Ren;J. Collomosse

文献摘要

被引文献

相似文献

我们描述了一个系统,用于在跨越近 100 年的大型跨媒体舞蹈录像档案中匹配人体姿势(姿势),其中包括排练和表演的数字照片和视频。由于其年代久远、质量和多样性,这段视频带来了独特的挑战。我们提出了一种类似森林的姿势表示,结合了多个尺度的视觉结构(自相似性)描述符,而没有显式检测肢体位置,这对于我们的数据来说是不可行的。我们探索两种互补的多尺度表示,将受文本检索领域启发的段落检索和潜在狄利克雷分配(LDA)技术应用于姿势匹配问题。结果是一个强大的系统,能够快速搜索大型跨媒体集合,以查找与视觉上指定的查询姿势的相似性。我们使用舞蹈专业人士提供的视觉查询来评估英国国家舞蹈研究中心 (UK-NRCD) 和 Siobhan Davies Replay (SDR) 数字舞蹈档案的横截面。我们展示了两个基线的显着性能改进:经典的单尺度和多尺度视觉词袋(BoVW)和空间金字塔内核(SPK)匹配。
We describe a system for matching human posture (pose) across a large cross-media archive of dance footage spanning nearly 100 years, comprising digitized photographs and videos of rehearsals and performances. This footage presents unique challenges due to its age, quality and diversity. We propose a forest-like pose representation combining visual structure (self-similarity) descriptors over multiple scales, without explicitly detecting limb positions which would be infeasible for our data. We explore two complementary multi-scale representations, applying passage retrieval and latent Dirichlet allocation (LDA) techniques inspired by the text retrieval domain, to the problem of pose matching. The result is a robust system capable of quickly searching large cross-media collections for similarity to a visually specified query pose. We evaluate over a cross-section of the UK National Research Centre for Dance's (UK-NRCD), and the Siobhan Davies Replay's (SDR) digital dance archives, using visual queries supplied by dance professionals. We demonstrate significant performance improvements over two base-lines: classical single and multi-scale bag of visual words (BoVW) and spatial pyramid kernel (SPK) matching .