HandyGaze: A Gaze Tracking Technique for Room-Scale Environments using a Single Smartphone

HandyGaze: A Gaze Tracking Technique for Room-Scale Environments using a Single Smartphone
复制标题

DOI:
10.1145/3567715
复制
发表时间:
2022-11
影响因子:
--
通讯作者:
Takahiro Nagai;Kazuyuki Fujita;Kazuki Takashima;Y. Kitamura
Takahiro Nagai;Kazuyuki Fujita;Kazuki Takashima;Y. Kitamura
中科院分区:
--
文献类型:
--
作者:
Takahiro Nagai;Kazuyuki Fujita;Kazuki Takashima;Y. Kitamura

文献摘要

相似文献

我们提出了 HandyGaze,这是一种适用于房间规模环境的 6 自由度注视跟踪技术,只需自然地握住智能手机即可进行,无需在环境中安装任何传感器或标记。我们的技术同时使用智能手机的前置和后置摄像头:前置摄像头估计用户相对于智能手机的注视矢量,而后置摄像头(和深度传感器,如果有)通过重建预先获得的环境 3D 地图来执行自定位。为了实现这一目标,我们通过运行基于 ARKit 的算法来估计用户的 6-DoF 头部方向,从而实现了一个适用于 iOS 智能手机的原型。我们还实现了一种新颖的校准方法,可以抵消用户特定的头部和注视方向之间的偏差。然后,我们进行了一项用户研究 (N=10),根据使用和不使用深度传感器和校准的组合,测量了我们的技术在四种条件下对注视目标的位置精度。结果表明,我们的校准方法能够将注视点的平均绝对误差降低 27%,使用深度传感器时误差为 0.53 m。我们还报告了避免错误输入所需的目标大小。最后,我们建议可能的应用,例如博物馆基于凝视的引导应用。
We propose HandyGaze, a 6-DoF gaze tracking technique for room-scale environments that can be carried out by simply holding a smartphone naturally without installing any sensors or markers in the environment. Our technique simultaneously employs the smartphone’s front and rear cameras: The front camera estimates the user’s gaze vector relative to the smartphone, while the rear camera (and depth sensor, if available) performs self-localization by reconstructing a pre-obtained 3D map of the environment. To achieve this, we implemented a prototype that works on iOS smartphones by running an ARKit-based algorithm for estimating the user’s 6-DoF head orientation. We additionally implemented a novel calibration method that offsets the user-specific deviation between the head and gaze orientations. We then conducted a user study (N=10) that measured our technique’s positional accuracy to the gaze target under four conditions, based on combinations of use with and without a depth sensor and calibration. The results show that our calibration method was able to reduce the mean absolute error of the gaze point by 27%, with an error of 0.53 m when using the depth sensor. We also report the target size required to avoid erroneous inputs. Finally, we suggest possible applications such as a gaze-based guidance application for museums.