"Binary" and "non-binary" detection tasks: Are current performance measures optimal?

"Binary" and "non-binary" detection tasks: Are current performance measures optimal?
复制标题

DOI:
10.1016/j.acra.2007.03.014
复制
发表时间:
2007-07-01
期刊:
影响因子:
4.8
通讯作者:
Bandos, Andriy I.
Bandos, Andriy I.
中科院分区:
医学3区
文献类型:
--
作者:
Gur, David;Rockette, Howard E.;Bandos, Andriy I.

文献摘要

被引文献

相似文献

基本原理和目标。我们观察到,在观察者研究的执行过程中,多项检测任务的很大一部分响应都处于低于 11% 或高于 89% 的极端范围,无论实际存在或不存在所讨论的异常或其主观评价的“微妙性”。这一观察结果提出了关于使用多类别评分量表进行此类检测任务的有效性和适当性的问题。对这些任务的二元和多类别评级的蒙特卡洛模拟表明,使用前者(二元)通常会产生偏差较小且更精确的汇总指数,因此可能会导致确定模态之间差异的统计功效更高。
Rationale and Objectives. We have observed that a very large fraction of responses for several detection tasks during the performance of observer studies are in the extreme ranges of lower than 11% or higher than 89% regardless of the actual presence or absence of the abnormality in question or its subjectively rated "subtleness." This observation raises questions regarding the validity and appropriateness of using multicategory rating scales for such detection tasks. Monte Carlo simulation of binary and multicategory ratings for these tasks demonstrate that the use of the former (binary) often results in a less biased and more precise summary index and hence may lead to a higher statistical power for determining differences between modalities.