Space-Time Correspondence as a Contrastive Random Walk

Space-Time Correspondence as a Contrastive Random Walk
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Jabri;Andrew Owens;Alexei A. Efros
A. Jabri;Andrew Owens;Alexei A. Efros
中科院分区:
其他
文献类型:
--
作者:
A. Jabri;Andrew Owens;Alexei A. Efros

文献摘要

被引文献

相似文献

本文提出了一种简单的自我监督方法,用于学习从原始视频中进行视觉对应的学习表示。我们将信件作为链接预测在视频构建的时空图中。在此图中,节点是从每个帧中采样的补丁,并且时间相邻的节点可以共享有向的边缘。我们学习一个节点嵌入,其中成对相似性定义了随机步行的过渡概率。远程对应关系的预测有效地计算为沿此图的步行。嵌入式学会通过沿对应的路径放置高概率来指导步行。目标是通过周期一致性而无需监督的:我们训练嵌入以最大程度地沿着从“ palindrome”框架构造的图时返回初始节点的可能性。我们证明该方法允许从大型未标记视频中学习表示。尽管它很简单,但该方法在涉及对象,语义部分和姿势的各种标签传播任务上的自我监督的最新范围都优于自我监督的最先进。此外,我们表明在测试时间和边缘辍学时进行自我监督的适应性改善了对象级对应的传输。
This paper proposes a simple self-supervised approach for learning representations for visual correspondence from raw video. We cast correspondence as link prediction in a space-time graph constructed from a video. In this graph, the nodes are patches sampled from each frame, and nodes adjacent in time can share a directed edge. We learn a node embedding in which pairwise similarity defines transition probabilities of a random walk. Prediction of long-range correspondence is efficiently computed as a walk along this graph. The embedding learns to guide the walk by placing high probability along paths of correspondence. Targets are formed without supervision, by cycle-consistency: we train the embedding to maximize the likelihood of returning to the initial node when walking along a graph constructed from a `palindrome' of frames. We demonstrate that the approach allows for learning representations from large unlabeled video. Despite its simplicity, the method outperforms the self-supervised state-of-the-art on a variety of label propagation tasks involving objects, semantic parts, and pose. Moreover, we show that self-supervised adaptation at test-time and edge dropout improve transfer for object-level correspondence.