Gaze Estimation Using Residual Neural Network

Gaze Estimation Using Residual Neural Network
复制标题

使用残差神经网络的注视估计

DOI:
10.1109/percomw.2019.8730846
复制
发表时间:
2019
期刊:
2019 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops)
影响因子:
--
通讯作者:
D. Rajan
D. Rajan
中科院分区:
--
文献类型:
--
作者:
En Teng Wong;Seanglidet Yean;Q. Hu;Bu;Jigang Liu;D. Rajan

文献摘要

被引文献

相似文献

视线跟踪已经成为人机交互和计算机视觉领域的一个重要研究课题。这是由于它在许多领域的应用,如市场调查,医学,神经科学和心理学。眼睛注视跟踪是通过估计离线或实时捕获的视频中的每个单独帧的注视(注视估计)来实现的。因此,为了产生安全准确的跟踪,特别是在医疗和社区的新兴应用中,视线估计的创新是研究领域的一个挑战。在本文中,我们探索了使用深度学习模型,残差神经网络(ResNet-18),来预测移动终端上的眼睛注视。该模型使用名为GazeCapture的大规模眼动跟踪公共数据集进行训练。我们的目标是通过结合去除闪烁数据的方法/技术,应用图像直方图归一化,头部姿势和面部网格特征进行创新。结果,我们实现了3.05厘米的平均误差,这比iTracker(4.11厘米的平均误差)更好,iTracker是最近使用AlexNet架构的凝视跟踪深度学习模型。经过观察,发现图像的自适应归一化比直方图归一化产生更好的结果。此外,我们发现头部姿势信息对所提出的深度学习网络是有用的贡献,而面部网格信息无助于减少测试误差。
Eye gaze tracking has become an prominent research topic in human-computer interaction and computer vision. It is due to its application in numerous fields, such as the market research, medical, neuroscience and psychology. Eye gaze tracking is implemented by estimating gaze (gaze estimation) for each individual frame in offline or real-time video captured. Therefore, in order to produce the secure the accurate tracking, especially in the emerging use in medical and community, innovation on the gaze estimation posts a challenge in research field. In this paper, we explored the use of the deep learning model, Residual Neural Network (ResNet-18), to predict the eye gaze on mobile device. The model is trained using the large-scale eye tracking public dataset called GazeCapture. We aim to innovate by incorporating methods/techniques of removing the blinking data, applying image histogram normalisation, head pose, and face grid features. As a result, we achieved 3.05cm average error, which is better performance than iTracker (4.11cm average error), the recent gaze tracking deep-learning model using AlexNet architecture. Upon observation, adaptive normalisation of the images was found to produce better results compared to histogram normalisation. Additionally, we found that head pose information was useful contribution to the proposed deep-learning network, while face grid information does not help to reduce test error.