DASZL: Dynamic Action Signatures for Zero-shot Learning

DASZL: Dynamic Action Signatures for Zero-shot Learning
复制标题

DOI:
10.1609/aaai.v35i3.16276
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
Tae Soo Kim;Jonathan D. Jones;Michael Peven;Zihao Xiao;Jin Bai;Yi Zhang;Weichao Qiu;A. Yuille-A.-Yuill
Tae Soo Kim;Jonathan D. Jones;Michael Peven;Zihao Xiao;Jin Bai;Yi Zhang;Weichao Qiu;A. Yuille-A.-Yuill
中科院分区:
其他
文献类型:
--
作者:
Tae Soo Kim;Jonathan D. Jones;Michael Peven;Zihao Xiao;Jin Bai;Yi Zhang;Weichao Qiu;A. Yuille-A.-Yuill

文献摘要

相似文献

活动识别有许多实际应用,其中潜在活动描述的集合组合起来很大。这使得识别系统的端到端监督训练变得不切实际,因为没有训练集实际上能够包含整个标签集。在本文中,我们提出了一种细粒度识别方法,将活动建模为动态动作签名的组合。这种组合方法使我们能够将细粒度识别重新构建为零样本活动识别,其中检测器由深度学习组件支持的简单第一原理状态机“动态”组成。我们在 Olympic Sports 和 UCF101 数据集上评估我们的方法,我们的模型在多个实验范式下建立了新的最先进技术。我们还扩展了这种方法,形成了一个独特的零样本联合分割和视频活动分类框架,并展示了在广泛使用的手术数据集上对复杂动作序列进行零样本解码的第一个结果。最后,我们展示了我们可以使用现成的物体检测器来识别完全从头开始的设置中的活动,而无需额外的训练。
There are many realistic applications of activity recognition where the set of potential activity descriptions is combinatorially large. This makes end-to-end supervised training of a recognition system impractical as no training set is practically able to encompass the entire label set. In this paper, we present an approach to fine-grained recognition that models activities as compositions of dynamic action signatures. This compositional approach allows us to reframe fine-grained recognition as zero-shot activity recognition, where a detector is composed "on the fly" from simple first-principles state machines supported by deep-learned components. We evaluate our method on the Olympic Sports and UCF101 datasets, where our model establishes a new state of the art under multiple experimental paradigms. We also extend this method to form a unique framework for zero-shot joint segmentation and classification of activities in video and demonstrate the first results in zero- shot decoding of complex action sequences on a widely-used surgical dataset. Lastly, we show that we can use off-the-shelf object detectors to recognize activities in completely de-novo settings with no additional training.