Imaging technology and practice assessment studies: importance of the baseline or reference performance level.

Imaging technology and practice assessment studies: importance of the baseline or reference performance level.
复制标题

成像技术和实践评估研究:基线或参考性能水平的重要性。

DOI:
10.1148/radiol.2471070822
复制
发表时间:
2008
期刊:
影响因子:
19.7
通讯作者:
Gur,David
Gur,David
中科院分区:
医学1区
文献类型:
--
作者:
Gur,David

文献摘要

被引文献

相似文献

David Gur, ScD Many years ago, when technology developments and practice decisions were made by manufacturers and the radiology community, these were heavily dependent upon the “wow factor,” which was based primarily on the subjective assessments of but a few “authorities” or “thought leaders” in the field; only limited science had been included in the deliberations. Since then, we have become more sophisticated and scientifically minded in our decision-making process. Objective retrospective and prospective studies are now routinely used for technology and practice assessments (1–3). Business decisions and clinical practice preferences are frequently made on the basis of the results of these studies. The regulatory process has also relied heavily on objectively ascertained data (1, 4, 5). Unfortunately, because of cost and complexity, inferences (ie, conclusions) are often based solely (or primarily) on the results of one “pivotal” study, possibly ignoring the fact that this very study may not be representative of actual clinical practice (6). Worse yet, the study results of the current practice, which serve as the baseline or reference performance level for comparison with the results of the new technology or practice being evaluated, may not reflect a generally accepted performance level. In addition, perspective regarding the measured reference performance level in these studies is often brief and frequently completely ignored. With increasing experience in assessing value, particularly in terms of demonstrated increases in efficiencies, effectiveness, and diagnostic accuracy or decreases in costs (or a combination thereof), the role of appropriate interpretation of scientific observations has changed dramatically. It is our responsibility to convey inferences that are generalizable and likely to withstand the test of time in a rapidly changing environment. Furthermore, because of the possible implications of inferences we make, we should do our best to address issues that were not as important (or pertinent) in the past. This editorial will address but one important issue related to inferences made as a result of pivotal studies that I believe has been largely ignored: namely, the performance level of the reference (often termed baseline or current) technology or practice. Several of the more basic concepts used in this paper have been recently described in detail by Wagner et al (7) in an outstanding review article on systems evaluations and performance assessment of diagnostic systems. In the present editorial, the term operating point is defined as the sensitivity and specificity level of an observer rating a set of cases as either positive or negative (a binary decision) for the abnormality (or task) in question. This point represents a level of performance for the observer under the conditions in which the cases were read, and it depends on different factors, including but not limited to the case set, the conditions, the “aggressiveness”(threshold for positive or negative) under which the set was read, and the actual proficiency of the observer. A good example of operating points marked in the domain (space) of sensitivity and 1 J specificity for different observers reading the same set of cases is provided by Wagner et al (7) in figure 1. Generally, it can be assumed that the operating point of an observer would be but one point on the individual’s or group’s performance curve. Performance curve is defined here as the estimated curve of performance for an individual observer, a group, or a technology (eg, computeraided detection [CAD] alone) in the sensitivity and specificity domain, representing what should be the different sensitivity and specificity levels under different …