Full receiver operating characteristic curve estimation using two alternative forced choice studies.

Full receiver operating characteristic curve estimation using two alternative forced choice studies.
复制标题

使用两个替代的强制选择研究来估计完整的受试者工作特征曲线。

DOI:
10.1117/1.jmi.3.1.011010
复制
发表时间:
2016
期刊:
Journal of medical imaging (Bellingham, Wash.)
影响因子:
--
通讯作者:
Brankov,JovanG
Brankov,JovanG
中科院分区:
--
文献类型:
--
作者:
Massanes,Francesc;Brankov,JovanG

文献摘要

被引文献

相似文献

基于任务的医学图像质量通常通过人类观察者在心理物理学人类观察者研究中执行诊断任务的程度来衡量。在典型的研究中,观察者被要求提供一个数字分数,量化他对图像是否包含诊断标记的信心。然后使用这些分数来衡量观察者的诊断准确性,并通过受试者工作特征 (ROC) 曲线和 ROC 曲线下面积进行总结。这些类型的人体研究很难安排、成本高昂且耗时。此外,参与此类研究的人类观察者应该是图像类型方面的专家,以避免在漫长的研究中得分不一致。在已知速度更快的两种替代强制选择 (2AFC) 研究中,同时比较两个图像并给出一个指标。不幸的是,2AFC 方法无法得出完整的 ROC 曲线或一组图像分数。这项工作的目的是提出一种方法,其中使用多轮 2AFC 研究来重新估计图像置信度得分(也称为评级、排名)并生成完整的 ROC 曲线。在所提出的方法中,我们将图像置信度得分视为需要估计的未知评分,并将 2AFC 视为两人比赛游戏。为了实现这一目标,我们使用 ELO 评级系统,该系统用于计算国际象棋等竞争者与竞争者游戏中玩家的相对技能水平。所提出的方法不限于 ELO,还可以使用其他评级方法,例如 TrueSkill™、Chessmetrics 或 Glicko。使用模拟数据得出的结果表明,可以使用多轮 2AFC 研究恢复完整的 ROC 曲线,并且最佳配对策略从第一轮配对异常图像与正常图像(如经典 2AFC 方法)开始,然后进行多轮随机配对。此外,所提出的方法在人类观察者试点研究中进行了测试。这些试点结果表明,三到五轮 2AFC 研究比完整评分研究需要更少的人类观察时间,并且重新估计的 ROC 曲线和 ROC 曲线下相关面积值与完整评分研究具有高度统计一致性。
Task-based medical image quality is typically measured by the degree to which a human observer can perform a diagnostic task in a psychophysical human observer study. During a typical study, an observer is asked to provide a numerical score quantifying his confidence as to whether an image contains a diagnostic marker or not. Such scores are then used to measure the observers’ diagnostic accuracy, summarized by the receiver operating characteristic (ROC) curve and the area under ROC curve. These types of human studies are difficult to arrange, costly, and time consuming. In addition, human observers involved in this type of study should be experts on the image genre to avoid inconsistent scoring through the lengthy study. In two-alternative forced choice (2AFC) studies, known to be faster, two images are compared simultaneously and a single indicator is given. Unfortunately, the 2AFC approach cannot lead to a full ROC curve or a set of image scores. The aim of this work is to propose a methodology in which multiple rounds of the 2AFC studies are used to re-estimate an image confidence score (a.k.a. rating, ranking) and generate the full ROC curve. In the proposed approach, we treat image confidence score as an unknown rating that needs to be estimated and 2AFC as a two-player match game. To achieve this, we use the ELO rating system, which is used for calculating the relative skill levels of players in competitor-versus-competitor games such as chess. The proposed methodology is not limited to ELO, and other rating methods such as TrueSkill™, Chessmetrics, or Glicko can be also used. The presented results, using simulated data, indicate that a full ROC curve can be recovered using several rounds of 2AFC studies and that the best pairing strategy starts with the first round of pairing abnormal versus normal images (as in the classical 2AFC approach) followed by a number of rounds using random pairing. In addition, the proposed method was tested in a pilot human observer study. These pilot results indicate that three to five rounds of 2AFC studies require less human observer time than a full scoring study and that the re-estimated ROC curves and associated area under ROC curve values have high statistical agreement with the full scoring study.