Performance of the RAND Appropriateness Criteria
Performance of the RAND Appropriateness Criteria
复制标题
兰德适当性标准的表现
DOI:
10.1177/0272989x03252312
复制
发表时间:
2003
影响因子:
3.6
通讯作者:
A. Garber
中科院分区:
文献类型:
--
作者:
A. Barnato;A. Garber
I the 1980s, a group of researchers at RAND and UCLA developed an appropriateness rating method for the Health Services Utilization Study. Their work set the direction for the formal assessment of the appropriateness of medical and surgical interventions. The RAND team convened a 9-member expert panel composed of academic and community physicians representing specialties relevant to the procedure under evaluation. They prepared a comprehensive review of the literature regarding the indications for and outcomes of the procedure for the panelists to read. Each physician then rated hundreds of clinical scenarios on a scale of 1 (most inappropriate) to 9 (most appropriate). Panelists worked toward consensus by using a modified Delphi method before re-ranking each clinical scenario. When the RAND group’s reports of high rates of inappropriate hospitalization and procedure use began appearing in such publications as the Journal of the American Medical Association, New England Journal of Medicine, and the Lancet, they drew a great deal of attention from policy makers. Their method has been particularly influential because “appropriateness” is a familiar word that seems to be a powerful criterion for determining whether a procedure should or should not have been performed. Furthermore, it conforms to health services researchers’ suspicions about the capricious nature of geographic variations in the rates with which procedures are performed. Yet the apparent familiarity of the term also predisposes the appropriateness metric to misinterpretation. Appropriateness evaluation is not comprehensive, because it ignores underuse—failure to perform a procedure when it is appropriate. Furthermore, because it is based on a panel rating of hypothetical scenarios, it is insensitive to patient preferences. It is based on the opinions of experts who rely in part upon their own memories, interests, and perspectives, in addition to the literature review that is prepared for their use. Most important, the appropriateness rating method can produce biased estimates of rates of inappropriate treatment. In a 1993 New England Journal of Medicine article, Charles Phelps explored this bias by drawing an analogy with diagnostic tests. He demonstrated that as long as the (true) prevalence of inappropriate treatment is less than 50% and the “test” is not 100% sensitive and specific, the estimated rate of inappropriate care usually overestimates the true rate (see Figure 1). Hence, one needs to know the sensitivity and specificity of the RAND method in order to estimate the size of the bias in the appropriateness ratings. Estimating sensitivity and specificity ordinarily requires a “gold standard” against which the test can be compared. Because there is no gold standard for determining appropriateness, Phelps proposed 2 approaches for estimating the rate of misclassification: 1) simultaneous application of independent tests and 2) assessment of changes in health of patients who were classified as appropriately or inappropriately treated. The comparison of 2 independent tests (expert panels) on the same population for the evaluation of bypass surgery and for coronary revascularization and hysterectomy by RAND researchers confirmed that significant misclassification error exists. In this issue of Medical Decision Making, Tobacman and colleagues report the results of a study that takes the 2nd approach. Theirs is the 1st study to assess changes in health status (visual acuity) among patients classified
影响因子:
158.5
作者:
Shekelle, PG;Kahan, JP;Park, RE
通讯作者:
Park, RE