Comparison of Four Subjective Methods for Image Quality Assessment

Comparison of Four Subjective Methods for Image Quality Assessment
复制标题

DOI:
10.1111/j.1467-8659.2012.03188.x
复制
发表时间:
2012-12-01
影响因子:
2.5
通讯作者:
Mantiuk, Radoslaw
Mantiuk, Radoslaw
中科院分区:
计算机科学4区
文献类型:
--
作者:
Mantiuk, Rafal K.;Tomaszewska, Anna;Mantiuk, Radoslaw

文献摘要

被引文献

相似文献

为了提供一个令人信服的证据,证明一种新的方法比最先进的方法更好,计算机图形项目通常伴随着用户研究,其中一组观察员对几种算法的结果进行排名或评级。这样的用户研究,被称为主观图像质量评估实验,可能非常耗时,并且不能保证产生结论性的结果。本文旨在帮助设计高效和严格的质量评估实验,并强调结果分析的关键方面。为了促进数据分析的良好标准,我们回顾了数据分析的主要方法,如建立置信区间,统计检验和回顾性功效分析。两种方法的可视化排名结果与有意义的信息的统计和实际意义进行了探讨。最后,我们比较了四种主要的主观质量评价方法:单刺激,双刺激,强迫选择两两比较和相似性判断。我们的结论是,强制选择成对比较方法的结果在最小的测量方差,从而产生最准确的结果。假设比较条件的数量适中,这种方法也是最省时的。
To provide a convincing proof that a new method is better than the state of the art, computer graphics projects are often accompanied by user studies, in which a group of observers rank or rate results of several algorithms. Such user studies, known as subjective image quality assessment experiments, can be very time-consuming and do not guarantee to produce conclusive results. This paper is intended to help design efficient and rigorous quality assessment experiments and emphasise the key aspects of the results analysis. To promote good standards of data analysis, we review the major methods for data analysis, such as establishing confidence intervals, statistical testing and retrospective power analysis. Two methods of visualising ranking results together with the meaningful information about the statistical and practical significance are explored. Finally, we compare four most prominent subjective quality assessment methods: single-stimulus, double-stimulus, forced-choice pairwise comparison and similarity judgements. We conclude that the forced-choice pairwise comparison method results in the smallest measurement variance and thus produces the most accurate results. This method is also the most time-efficient, assuming a moderate number of compared conditions.