Fusing eye movements and observer narratives for expert-driven image-region annotations

Fusing eye movements and observer narratives for expert-driven image-region annotations
复制标题

融合眼球运动和观察者叙述,以进行专家驱动的图像区域注释

DOI:
10.1145/2857491.2857542
复制
发表时间:
2016
期刊:
Proceedings of the Ninth Biennial ACM Symposium on Eye Tracking Research & Applications
影响因子:
--
通讯作者:
Anne R. Haake
Anne R. Haake
中科院分区:
--
文献类型:
--
作者:
Preethi Vaidyanathan;J. Pelz;Emily Tucker Prud'hommeaux;Cecilia Ovesdotter Alm;Anne R. Haake

文献摘要

被引文献

相似文献

人类图像理解是通过个体的视觉和语言行为来反映的,但对个体的多模态表征进行有意义的计算整合和解释仍然是一个挑战。在本文中,我们扩展了一个框架,用于捕获图像区域的注释在皮肤病学,一个域中,解释图像的影响专家的视觉感知技能,概念领域的知识,和面向任务的目标。我们的工作探讨的假设,眼动可以帮助我们理解专家的知觉过程,口语描述可以揭示图像检查任务的概念元素。我们把有意义地整合视觉和语言数据的问题作为无监督的双文本对齐。使用对齐,我们在医生的眼球运动(揭示图像的关键区域)和这些图像的口头描述之间创建了有意义的映射。然后使用所得到的对准来用医学概念标签注释图像区域。我们的对齐精度超过基线使用精确和延迟的时间对应。此外,基于眼睛运动识别图像中的聚类的方法与使用图像特征识别聚类的方法之间的对准精度的比较表明,这两种方法在不同类型的图像和概念标签上表现良好。这表明图像注释框架应该集成来自多个技术的信息来处理异构图像。我们还研究了皮肤病学主要形态概念标签,以及病变大小或类型和分布为基础的类别的图像所提出的对准器的性能。
Human image understanding is reflected by individuals' visual and linguistic behaviors, but the meaningful computational integration and interpretation of their multimodal representations remain a challenge. In this paper, we expand a framework for capturing image-region annotations in dermatology, a domain in which interpreting an image is influenced by experts' visual perception skills, conceptual domain knowledge, and task-oriented goals. Our work explores the hypothesis that eye movements can help us understand experts' perceptual processes and that spoken language descriptions can reveal conceptual elements of image inspection tasks. We cast the problem of meaningfully integrating visual and linguistic data as unsupervised bitext alignment. Using alignment, we create meaningful mappings between physicians' eye movements, which reveal key areas of images, and spoken descriptions of those images. The resulting alignments are then used to annotate image regions with medical concept labels. Our alignment accuracy exceeds baselines using both exact and delayed temporal correspondence. Additionally, comparison of alignment accuracy between a method that identifies clusters in the images based on eye movement vs. a method that identifies clusters using image features suggests that the two approaches perform well on different types of images and concept labels. This suggests that an image annotation framework should integrate information from more than one technique to handle heterogeneous images. We also investigate the performance of the proposed aligner for dermatological primary morphology concept labels, as well as for lesion size or type and distribution-based categories of images.