OCTOPUS: Open-vocabulary Content Tracking and Object Placement Using Semantic Understanding in Mixed Reality

OCTOPUS: Open-vocabulary Content Tracking and Object Placement Using Semantic Understanding in Mixed Reality
复制标题

DOI:
10.1109/ismar-adjunct60411.2023.00124
复制
发表时间:
2023-10
期刊:
2023 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct)
影响因子:
--
通讯作者:
Luke Yoffe;Aditya Sharma;Tobias Höllerer
Luke Yoffe;Aditya Sharma;Tobias Höllerer
中科院分区:
其他
文献类型:
--
作者:
Luke Yoffe;Aditya Sharma;Tobias Höllerer

文献摘要

相似文献

增强现实的一个关键挑战是将虚拟内容放置在自然位置。现有的自动化技术只能处理封闭词汇表、固定的对象集。在本文中,我们介绍了一种新的开放词汇表的方法,对象放置。我们的八阶段流水线利用分割模型,视觉语言模型和LLM的最新进展,将任何虚拟对象放置在任何AR相机帧或场景中。在初步的用户研究中,我们表明,我们的方法至少在57%的时间内与人类专家一样好。1
One key challenge in augmented reality is the placement of virtual content in natural locations. Existing automated techniques are only able to work with a closed-vocabulary, fixed set of objects. In this paper, we introduce a new open-vocabulary method for object placement. Our eight-stage pipeline leverages recent advances in segmentation models, vision-language models, and LLMs to place any virtual object in any AR camera frame or scene. In a preliminary user study, we show that our method performs at least as well as human experts 57% of the time. 1