TokyoTechCanon at TRECVID 2012

TokyoTechCanon at TRECVID 2012
复制标题

TokyoTechCanon 参加 TRECVID 2012

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
K. Shinoda
K. Shinoda
中科院分区:
--
文献类型:
--
作者:
Nakamasa Inoue;Yusuke Kamishima;Kotaro Mori;K. Shinoda

文献摘要

参考文献

被引文献

相似文献

我们的目标是使用高斯混合模型 (GMM) 超向量和树结构 GMM [1,2,3] 开发高性能语义索引系统。从视频镜头中提取对应于六种音频和视觉特征的 GMM 超向量。树结构 GMM 降低了估计 GMM 参数的最大后验 (MAP) 适应的计算成本,同时保持高水平的准确性。今年,我们引入了 HOG-Dense 和 LBP-Dense 两个新的低级特征以及视频剪辑分数。 HOG-Dense 和 LBP-Dense 通过使用密集采样从每个镜头最多 100 帧中提取。视频剪辑得分被定义为视频剪辑中所有镜头中镜头得分的最大值,用于对视频镜头进行重新排名。我们的最佳结果是平均 InfAP 为 32.10%,在完整任务中的所有语义索引运行中排名第一。
We aim at developing a high-performance semantic indexing system using Gaussian-mixture-model (GMM) supervectors and tree-structured GMMs [1, 2, 3]. GMM supervectors corresponding to six types of audio and visual features are extracted from video shots. Tree-structured GMMs reduce the computational cost of maximum a posteriori (MAP) adaptation for estimating GMM parameters while keeping accuracy at high levels. This year, we introduce two new low-level features of HOG-Dense and LBP-Dense and video-clip scores. HOG-Dense and LBP-Dense are extracted from up to 100 frames per shot by using dense sampling. The video-clip score is defined as the maximum value of shot scores among all the shots in a video clip and is used for re-ranking video shots. Our best result was 32.10% in terms of Mean InfAP, which was ranked first over all semantic indexing runs in the full task.
DOI: 10.1007/11744023_32
发表时间: 2006-01-01
期刊: COMPUTER VISION - ECCV 2006 , PT 1, PROCEEDINGS
影响因子: --
作者:
Bay, Herbert;Tuytelaars, Tinne;Van Gool, Luc
通讯作者: Van Gool, Luc