Monocular Depth Estimation for Soft Visuotactile Sensors

Monocular Depth Estimation for Soft Visuotactile Sensors
复制标题

DOI:
10.1109/robosoft51838.2021.9479234
复制
发表时间:
2021-01
期刊:
2021 IEEE 4th International Conference on Soft Robotics (RoboSoft)
影响因子:
--
通讯作者:
Rares Ambrus;V. Guizilini;N. Kuppuswamy;Andrew Beaulieu;Adrien Gaidon;A. Alspach
Rares Ambrus;V. Guizilini;N. Kuppuswamy;Andrew Beaulieu;Adrien Gaidon;A. Alspach
中科院分区:
其他
文献类型:
--
作者:
Rares Ambrus;V. Guizilini;N. Kuppuswamy;Andrew Beaulieu;Adrien Gaidon;A. Alspach

文献摘要

被引文献

相似文献

填充液体的软性视觉触觉传感器(如Soft-bubbles)缓解了稳健操作的关键挑战,因为它们能够实现可靠的抓握沿着能够获得关于接触几何形状和力的高分辨率感官反馈。虽然它们结构简单,但由于封闭的定制IR/深度成像传感器直接测量表面变形所引入的尺寸限制,它们的实用性受到限制。为了减轻这一限制,我们研究了应用最先进的单目深度估计来直接从内部单个小型IR成像传感器推断密集的内部(触觉)深度图。通过真实世界的实验,我们表明,通常用于长距离深度估计(1- 100米)的深度网络可以有效地训练,以便在一个几乎无纹理的可变形流体填充传感器内进行更短距离(1- 100毫米)的精确预测。我们提出了一个简单的监督学习过程来训练一个对象不可知的网络,该网络需要少于10个随机姿势,接触不到10秒,用于一小组不同的对象(在我们的实验中,马克杯,酒杯,盒子和手指)。我们表明,我们的方法是样本高效,准确,并概括了不同的对象和传感器配置在训练时看不见。最后,我们讨论了我们的软视觉触觉传感器和夹具的设计方法的影响。
Fluid-filled soft visuotactile sensors such as the Soft-bubbles alleviate key challenges for robust manipulation, as they enable reliable grasps along with the ability to obtain high-resolution sensory feedback on contact geometry and forces. Although they are simple in construction, their utility has been limited due to size constraints introduced by enclosed custom IR/depth imaging sensors to directly measure surface deformations. Towards mitigating this limitation, we investigate the application of state-of-the-art monocular depth estimation to infer dense internal (tactile) depth maps directly from the internal single small IR imaging sensor. Through real-world experiments, we show that deep networks typically used for long-range depth estimation (1-100m) can be effectively trained for precise predictions at a much shorter range (1-100mm) inside a mostly textureless deformable fluid-filled sensor. We propose a simple supervised learning process to train an object-agnostic network requiring less than 10 random poses in contact for less than 10 seconds for a small set of diverse objects (mug, wine glass, box, and fingers in our experiments). We show that our approach is sample-efficient, accurate, and generalizes across different objects and sensor configurations unseen at training time. Finally, we discuss the implications of our approach for the design of soft visuotactile sensors and grippers1.