What Do Different Evaluation Metrics Tell Us About Saliency Models?

What Do Different Evaluation Metrics Tell Us About Saliency Models?
复制标题

DOI:
10.1109/tpami.2018.2815601
复制
发表时间:
2019-03-01
影响因子:
23.6
通讯作者:
Durand, Fredo
Durand, Fredo
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bylinskii, Zoya;Judd, Tilke;Durand, Fredo

文献摘要

被引文献

相似文献

如何最好地评估显着性模型预测人类在图像中的位置的能力是一个开放的研究问题。评估度量的选择取决于显著性如何定义以及地面真实如何表示。不同的显着性模型的排名方式不同,这取决于如何处理假阳性和假阴性,是否考虑了观看偏差,是否考虑了空间偏差,以及显着性图如何预处理。在本文中,我们提供了8个不同的评价指标和他们的属性进行了分析。借助系统的实验和可视化的度量计算,我们增加了显着性分数的可解释性和显着性模型评估的透明度。在度量属性和行为的差异的基础上,我们提出了在特定假设和特定应用下的度量选择建议。
How best to evaluate a saliency model's ability to predict where humans look in images is an open research question. The choice of evaluation metric depends on how saliency is defined and how the ground truth is represented. Metrics differ in how they rank saliency models, and this results from how false positives and false negatives are treated, whether viewing biases are accounted for, whether spatial deviations are factored in, and how the saliency maps are pre-processed. In this paper, we provide an analysis of 8 different evaluation metrics and their properties. With the help of systematic experiments and visualizations of metric computations, we add interpretability to saliency scores and more transparency to the evaluation of saliency models. Building off the differences in metric properties and behaviors, we make recommendations for metric selections under specific assumptions and for specific applications.