ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations

ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Ruohan Gao;Yen-Yu Chang;Shivani Mall;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu
Ruohan Gao;Yen-Yu Chang;Shivani Mall;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu
中科院分区:
其他
文献类型:
--
作者:
Ruohan Gao;Yen-Yu Chang;Shivani Mall;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu

文献摘要

被引文献

相似文献

多感官以物体为中心的感知、推理和交互是近年来的一个关键研究课题。然而,这些方向的进展受到可用物体数量少的限制——合成物体不够逼真,且大多围绕几何形状,而像YCB这样的真实物体数据集由于国际运输、库存和财务成本等原因,在获取上往往具有实际困难且不稳定。我们提出了ObjectFolder,这是一个包含100个虚拟物体的数据集,它通过两项关键创新解决了这两个挑战。首先,ObjectFolder对所有物体的视觉、听觉和触觉感知数据进行编码,实现了许多多感官物体识别任务,这超越了现有的仅仅关注物体几何形状的数据集。其次,ObjectFolder对每个物体的视觉纹理、声学模拟和触觉读数采用统一的、以物体为中心的隐式表示,使得该数据集使用灵活且易于共享。我们通过在各种基准任务(包括实例识别、跨感官检索、3D重建和机器人抓取)上对其进行评估,证明了我们的数据集作为多感官感知和控制测试平台的有用性。
Multisensory object-centric perception, reasoning, and interaction have been a key research topic in recent years. However, the progress in these directions is limited by the small set of objects available -- synthetic objects are not realistic enough and are mostly centered around geometry, while real object datasets such as YCB are often practically challenging and unstable to acquire due to international shipping, inventory, and financial cost. We present ObjectFolder, a dataset of 100 virtualized objects that addresses both challenges with two key innovations. First, ObjectFolder encodes the visual, auditory, and tactile sensory data for all objects, enabling a number of multisensory object recognition tasks, beyond existing datasets that focus purely on object geometry. Second, ObjectFolder employs a uniform, object-centric, and implicit representation for each object's visual textures, acoustic simulations, and tactile readings, making the dataset flexible to use and easy to share. We demonstrate the usefulness of our dataset as a testbed for multisensory perception and control by evaluating it on a variety of benchmark tasks, including instance recognition, cross-sensory retrieval, 3D reconstruction, and robotic grasping.