Binary and multi-category ratings in a laboratory observer performance study: a comparison.

Binary and multi-category ratings in a laboratory observer performance study: a comparison.
复制标题

实验室观察员绩效研究中的二元和多类别评级:比较。

DOI:
10.1118/1.2977766
复制
发表时间:
2008
期刊:
影响因子:
3.8
通讯作者:
Rockette,HowardE
Rockette,HowardE
中科院分区:
医学3区
文献类型:
--
作者:
Gur,David;Bandos,AndriyI;King,JillL;Klym,AmyH;Cohen,CathyS;Hakim,ChristianeM;Hardesty,LaraA;Ganott,MarieA;Perrin,RonaldL;Poller,WilliamR;Shah,Ratan;Sumkin,JulesH;Wallace,LuisaP;Rockette,HowardE

文献摘要

相似文献

作者调查了放射科医生在回顾性解读筛查乳腺X线照片期间的表现,当使用二元决策是否召回女性进行额外手术时,并使用半连续评级量表将其与其受试者工作特征(ROC)型性能曲线进行比较。根据机构审查委员会批准的方案,9名经验丰富的放射科医生使用筛选BI-RADS评定量表(回忆/不回忆)和半连续ROC类型评定量表(0至100),对他们未在诊所亲自阅读的155项检查的丰富集进行独立评定,并与他们在诊所单独阅读的其他丰富集混合。计算每个阅片者的经验ROC曲线和二元操作点之间的垂直距离,即相同特异性水平下的灵敏度水平差异。使用所有阅片员的平均垂直距离评估二元和ROC型评定量表下性能水平的接近程度。当使用两种评定方法中的任何一种时,读片者似乎没有任何系统性的倾向于更好的表现,即四名读片者使用半连续评定量表表现更好,四名读片者使用二进制量表表现更好,一名读片者的点正好在经验ROC曲线上。九个读者中只有一个有一个二进制的“工作点”,这是统计上远离同一读者的经验ROC曲线。读片员特异性差异范围为0.128,相应95%置信区间的平均宽度为0.2,个体读片员的数值范围为0.050 - 0.966。平均而言,放射科医生在使用两个评级量表时表现相似,因为个体阅片者的二元操作点与其ROC曲线之间的平均距离接近于零。固定阅片人平均值(0.016)的95%置信区间为(,0.0631)(双侧值0.35)。总之,作者发现,在回顾性观察者性能研究中,使用二元响应或半连续评定量表可在通过灵敏度-特异性操作点测量的性能方面获得一致的结果。
The authors investigated radiologists, performances during retrospective interpretation of screening mammograms when using a binary decision whether to recall a woman for additional procedures or not and compared it with their receiver operating characteristic (ROC) type performance curves using a semi‐continuous rating scale. Under an Institutional Review Board approved protocol nine experienced radiologists independently rated an enriched set of 155 examinations that they had not personally read in the clinic, mixed with other enriched sets of examinations that they had individually read in the clinic, using both a screening BI‐RADS rating scale (recall/not recall) and a semi‐continuous ROC type rating scale (0 to 100). The vertical distance, namely the difference in sensitivity levels at the same specificity levels, between the empirical ROC curve and the binary operating point were computed for each reader. The vertical distance averaged over all readers was used to assess the proximity of the performance levels under the binary and ROC‐type rating scale. There does not appear to be any systematic tendency of the readers towards a better performance when using either of the two rating approaches, namely four readers performed better using the semi‐continuous rating scale, four readers performed better with the binary scale, and one reader had the point exactly on the empirical ROC curve. Only one of the nine readers had a binary “operating point” that was statistically distant from the same reader's empirical ROC curve. Reader‐specific differences ranged from to 0.128 with an average width of the corresponding 95% confidence intervals of 0.2 and ‐values ranging for individual readers from 0.050 to 0.966. On average, radiologists performed similarly when using the two rating scales in that the average distance between the run in individual reader's binary operating point and their ROC curve was close to zero. The 95% confidence interval for the fixed‐reader average (0.016) was (, 0.0631) (two‐sided ‐value 0.35). In conclusion the authors found that in retrospective observer performance studies the use of a binary response or a semi‐continuous rating scale led to consistent results in terms of performance as measured by sensitivity‐specificity operating points.