Data-Driven, 3-D Classification of Person-Object Relationships and Semantic Context Clustering for Robotics and AI Applications

Data-Driven, 3-D Classification of Person-Object Relationships and Semantic Context Clustering for Robotics and AI Applications
复制标题

DOI:
10.1109/roman.2018.8525654
复制
发表时间:
2018-08
期刊:
2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)
影响因子:
--
通讯作者:
M. P. Zapf;Astha Gupta;Luis Yoichi Morales Saiki;M. Kawanabe
M. P. Zapf;Astha Gupta;Luis Yoichi Morales Saiki;M. Kawanabe
中科院分区:
其他
文献类型:
--
作者:
M. P. Zapf;Astha Gupta;Luis Yoichi Morales Saiki;M. Kawanabe

文献摘要

相似文献

我们介绍了一个框架,用于检测和分类的时空人与物体的相互作用。我们的方法从RGB-D数据检测到的交互中聚类出相似的语义上下文。2-D物体检测(YOLO)基于来自移动的机器人上的Kinect v2传感器的RGB数据运行,该机器人在办公室中导航并观察人员和办公桌空间。通过RGB深度配准和连续的欧几里德和k均值空间聚类将人和物体检测转换为3-D点云时间序列。3-D人和物体点云流用于创建时间序列占用图和人-物体协同定位图。从这些地图,人和不同的对象之间的时空相关性计算。使用k-均值聚类相关模式,以获得不同的人-对象交互,即随着时间的推移段语义上下文。我们通过记录90个30秒的RGB-D数据集来评估我们的方法检测人-物相关性和聚类语义上下文的性能,其中三个人处理代表性对象(书籍,杯子,瓶子)。实验结果表明,我们的框架是能够始终如一地分配语义上下文相同的集群在> 79%的情况下(场景帧)。视觉场景中的语义上下文可以区分,而不需要提供先验信息,允许移动的代理学习和探索新的环境。
We introduce a framework for detection and classification of spatio-temporal person-object interactions. Our method clusters similar semantic contexts from interactions detected from RGB-D data. 2-D object detection (YOLO) is run on RGB data from a Kinect v2 sensor on a mobile robot navigating an office and observing persons and desk spaces. Person and object detections are converted into 3-D point cloud time series via RGB-Depth co-registration and successive Euclidean and k-means spatial clustering. 3-D person and object point cloud streams are used to create time-series occupancy maps and person-object co-localization maps. From these maps, spatiotemporal correlations between persons and distinct objects are computed. Correlation patterns are clustered using k-means to obtain distinct human-object interactions, i.e. segment semantic context over time. We evaluated the performance of our approach to detect person-object correlations and cluster semantic context by recording 90 30-second RGB-D data episodes, with three persons handling representative objects (books, cups, bottles). Experimental results show that our framework is able to consistently assign semantic context to the same cluster in > 79% of cases (scene frames). Semantic contexts in visual scenes can be distinguished without the need to provide prior information, allowing mobile agents to learn and explore in new environments.