Performance of the RAND Appropriateness Criteria

Performance of the RAND Appropriateness Criteria
复制标题

兰德适当性标准的表现

DOI:
10.1177/0272989x03252312
复制
发表时间:
2003
影响因子:
3.6
通讯作者:
A. Garber
A. Garber
中科院分区:
医学3区
文献类型:
--
作者:
A. Barnato;A. Garber

文献摘要

参考文献

被引文献

相似文献

20世纪80年代,兰德公司和加州大学洛杉矶分校的一组研究人员为卫生服务利用研究开发了一种适当性评级方法。他们的工作为正式评估医疗和外科干预措施的适当性确定了方向。兰德团队召集了一个由学术和社区医生组成的9人专家小组,代表与评估过程相关的专业。他们准备了一份关于手术适应症和结果的文献综述,供小组成员阅读。然后,每位医生对数百种临床情况按1(最不合适)到9(最合适)的等级进行评分。在对每个临床情景重新排序之前,小组成员通过使用改进的德尔菲法达成共识。当兰德集团关于不适当住院和手术使用的高比率的报告开始出现在诸如《美国医学会杂志》、《新英格兰医学杂志》和《柳叶刀》等出版物上时,它们引起了政策制定者的极大关注。他们的方法特别有影响力,因为“适当性”是一个熟悉的词,似乎是决定一个程序是否应该进行的有力标准。此外,它符合卫生服务研究人员的怀疑,即手术执行率的地理差异是反复无常的。然而,对该术语的明显熟悉也使适当性度量容易被误解。适当性评估是不全面的,因为它忽略了未充分利用——在适当的时候没有执行程序。此外,由于它是基于假设情景的小组评级,因此对患者的偏好不敏感。它基于专家的意见,这些专家部分依赖于他们自己的记忆、兴趣和观点,以及为他们的使用准备的文献综述。最重要的是,适当性评级方法可能产生不适当治疗率的有偏差估计。在1993年《新英格兰医学杂志》上的一篇文章中,查尔斯·菲尔普斯(Charles Phelps)通过与诊断测试进行类比,探讨了这种偏见。他证明,只要不适当治疗的(真实)患病率低于50%,并且“测试”不是100%敏感和特异性,则估计的不适当治疗率通常高估了真实率(见图1)。因此,人们需要知道RAND方法的敏感性和特异性,以便估计适当性评级中的偏差大小。估计灵敏度和特异性通常需要一个“金标准”,以便与测试进行比较。由于没有确定适当性的金标准,Phelps提出了两种估计错误分类率的方法:1)同时应用独立测试;2)评估被分类为适当或不适当治疗的患者的健康变化。RAND研究人员对同一人群进行的评估搭桥手术和冠状动脉血运重建术及子宫切除术的2个独立试验(专家小组)的比较证实存在显著的误分类错误。在这一期的《医疗决策》中,Tobacman和他的同事报告了采用第二种方法的研究结果。他们的研究是第一个评估患者健康状况(视力)变化的研究
I the 1980s, a group of researchers at RAND and UCLA developed an appropriateness rating method for the Health Services Utilization Study. Their work set the direction for the formal assessment of the appropriateness of medical and surgical interventions. The RAND team convened a 9-member expert panel composed of academic and community physicians representing specialties relevant to the procedure under evaluation. They prepared a comprehensive review of the literature regarding the indications for and outcomes of the procedure for the panelists to read. Each physician then rated hundreds of clinical scenarios on a scale of 1 (most inappropriate) to 9 (most appropriate). Panelists worked toward consensus by using a modified Delphi method before re-ranking each clinical scenario. When the RAND group’s reports of high rates of inappropriate hospitalization and procedure use began appearing in such publications as the Journal of the American Medical Association, New England Journal of Medicine, and the Lancet, they drew a great deal of attention from policy makers. Their method has been particularly influential because “appropriateness” is a familiar word that seems to be a powerful criterion for determining whether a procedure should or should not have been performed. Furthermore, it conforms to health services researchers’ suspicions about the capricious nature of geographic variations in the rates with which procedures are performed. Yet the apparent familiarity of the term also predisposes the appropriateness metric to misinterpretation. Appropriateness evaluation is not comprehensive, because it ignores underuse—failure to perform a procedure when it is appropriate. Furthermore, because it is based on a panel rating of hypothetical scenarios, it is insensitive to patient preferences. It is based on the opinions of experts who rely in part upon their own memories, interests, and perspectives, in addition to the literature review that is prepared for their use. Most important, the appropriateness rating method can produce biased estimates of rates of inappropriate treatment. In a 1993 New England Journal of Medicine article, Charles Phelps explored this bias by drawing an analogy with diagnostic tests. He demonstrated that as long as the (true) prevalence of inappropriate treatment is less than 50% and the “test” is not 100% sensitive and specific, the estimated rate of inappropriate care usually overestimates the true rate (see Figure 1). Hence, one needs to know the sensitivity and specificity of the RAND method in order to estimate the size of the bias in the appropriateness ratings. Estimating sensitivity and specificity ordinarily requires a “gold standard” against which the test can be compared. Because there is no gold standard for determining appropriateness, Phelps proposed 2 approaches for estimating the rate of misclassification: 1) simultaneous application of independent tests and 2) assessment of changes in health of patients who were classified as appropriately or inappropriately treated. The comparison of 2 independent tests (expert panels) on the same population for the evaluation of bypass surgery and for coronary revascularization and hysterectomy by RAND researchers confirmed that significant misclassification error exists. In this issue of Medical Decision Making, Tobacman and colleagues report the results of a study that takes the 2nd approach. Theirs is the 1st study to assess changes in health status (visual acuity) among patients classified
DOI: 10.1056/nejm199806253382607
发表时间: 1998-06-25
影响因子: 158.5
作者:
Shekelle, PG;Kahan, JP;Park, RE
通讯作者: Park, RE