Convolutional Neural Network-Based Visual Servoing for Eye-to-Hand Manipulator

Convolutional Neural Network-Based Visual Servoing for Eye-to-Hand Manipulator
复制标题

DOI:
10.1109/access.2021.3091737
复制
发表时间:
2021-01-01
期刊:
影响因子:
3.9
通讯作者:
Kosuge, Kazuhiro
Kosuge, Kazuhiro
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tokuda, Fuyuki;Arai, Shogo;Kosuge, Kazuhiro

文献摘要

被引文献

相似文献

我们提出了一种基于CNN的视觉伺服方案,用于眼手机械手的精确定位,其中机器人的控制输入通过神经网络直接从图像中计算出来。本文提出了一种新的卷积神经网络--编码特征驱动差分交互矩阵网络(DEFINet),用于眼手视觉伺服。DEFINet从目视相机捕获的期望和当前图像中估计期望和当前终端效应器之间的相对姿势。DEFINet包括同一个CNN的两个分支,它们共享权重并对目标图像和当前图像进行编码,这是受到暹罗网络架构的启发。利用编码后的目标和当前图像特征的差异回归相对位姿,可以获得较高的DEFINet视觉伺服定位精度。训练数据集是通过在任务空间中随机操作机械手而收集的样本数据生成的。通过数值仿真和在真实环境中使用六自由度工业机械手的实验,对所提出的视觉伺服的性能进行了评估。仿真和实验结果表明了该方法的有效性。
We propose a CNN based visual servoing scheme for precise positioning of an eye-to-hand manipulator in which the control input of a robot is calculated directly from images by a neural network. In this paper, we propose Difference of Encoded Features driven Interaction matrix Network (DEFINet), a new convolutional neural network (CNN), for eye-to-hand visual servoing. DEFINet estimates a relative pose between desired and current end-effector from desired and current images captured by an eye-to-hand camera. DEFINet includes two branches of the same CNN that share weights and encode target and current images, which is inspired by the architecture of Siamese network. Regression of the relative pose from the difference of the encoded target and current image features leads to a high positioning accuracy of visual servoing using DEFINet. The training dataset is generated from sample data collected by operating a manipulator randomly in task space. The performance of the proposed visual servoing is evaluated through numerical simulation and experiments using a six-DOF industrial manipulator in a real environment. Both simulation and experimental results show the effectiveness of the proposed method.