A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs

A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
复制标题

DOI:
10.1109/icra48506.2021.9561271
复制
发表时间:
2020-11
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Juntao Tan;Changkyu Song;Abdeslam Boularias
Juntao Tan;Changkyu Song;Abdeslam Boularias
中科院分区:
其他
文献类型:
--
作者:
Juntao Tan;Changkyu Song;Abdeslam Boularias

文献摘要

相似文献

本文提出了一种新的自监督系统,用于学习检测图像中新的和以前未见的对象类别。所提出的系统接收包含各种对象的场景的几个未标记视频作为输入。使用深度信息将视频的帧分割成对象,并沿着每个视频跟踪片段。然后,该系统构建一个加权图,该图基于序列所包含的对象之间的相似性来连接序列。在自动重新排列两个序列中的帧以对准对象的视点之后,通过使用通用视觉特征来测量两个对象序列之间的相似性。该图用于通过执行随机遍历对相似和不相似示例的三元组进行采样。最后使用三元组样本来训练暹罗神经网络,该神经网络将一般视觉特征投影到低维流形中。在YCB-Video、CORe50和RGBD-Object三个公共数据集上的实验表明,投影的低维特征提高了未知对象到新类别的聚类精度,并优于最近几种非监督聚类方法。
This paper presents a new self-supervised system for learning to detect novel and previously unseen categories of objects in images. The proposed system receives as input several unlabeled videos of scenes containing various objects. The frames of the videos are segmented into objects using depth information, and the segments are tracked along each video. The system then constructs a weighted graph that connects sequences based on the similarities between the objects that they contain. The similarity between two sequences of objects is measured by using generic visual features, after automatically re-arranging the frames in the two sequences to align the viewpoints of the objects. The graph is used to sample triplets of similar and dissimilar examples by performing random walks. The triplet examples are finally used to train a siamese neural network that projects the generic visual features into a low-dimensional manifold. Experiments on three public datasets, YCB-Video, CORe50 and RGBD-Object, show that the projected low-dimensional features improve the accuracy of clustering unknown objects into novel categories, and outperform several recent unsupervised clustering techniques.