Reliability of the PEDro scale for rating quality of randomized controlled trials

Reliability of the PEDro scale for rating quality of randomized controlled trials
复制标题

DOI:
10.1093/ptj/83.8.713
复制
发表时间:
2003-08-01
期刊:
影响因子:
3.2
通讯作者:
Elkins, M
Elkins, M
中科院分区:
医学2区
文献类型:
--
作者:
Maher, CG;Sherrington, C;Elkins, M

文献摘要

被引文献

相似文献

背景和目的。随机对照试验(RCT)的质量评估是系统评价中的常见做法。然而,用大多数质量评估表获得的数据的可靠性尚未确定。这份报告描述了两项研究,旨在调查使用物理治疗证据数据库(PEDRO)量表获得的数据的可靠性,该量表的开发是为了评估物理治疗师干预的随机对照试验的质量。方法。在第一项研究中,11名评分员从PEDRO数据库中随机挑选了25名随机对照试验,独立对其进行评分。在第二项研究中,两名评分员从PEDRO数据库中随机挑选了120名随机对照试验进行评分,分歧由第三名评分员解决;这产生了一组个人评分员和共识评分。独立评分者重复了这一过程,创建了第二套个人和共识评级。使用多评分者Kappas计算Pedro量表条目评分的可靠性,并使用组内相关系数(ICC[1,1])计算总分(总分)的可靠性。结果。每个I-I项目的kappa值在0.36到0.80之间,对于由2或3个评分者组成的团体产生的共识评分在0.50到0.79之间。对于个人评分,总分的ICC为0.56(95%可信区间=0.47-0.65),对于共识评分,ICC为0.68(95%可信区间=0.57-0.76)。讨论和结论。佩德罗量表项目的可信度从“一般”到“相当”不等,佩德罗总分的可信度从“一般”到“良好”。
Background and Purpose. Assessment of the quality of randomized controlled trials (RCTs) is common practice in systematic reviews. However; the reliability of data obtained with most quality assessment scales has not been established. This report describes 2 studies designed to investigate the reliability of data obtained with the Physiotherapy Evidence Database (PEDro) scale developed to rate the quality of RCTs evaluating physical therapist interventions. Method. In the first study, 11 raters independently rated 25 RCTs randomly selected from the PEDro database. In the second study, 2 raters rated 120 RCTs randomly selected from the PEDro database, and disagreements were resolved by a third rater; this generated a set of individual rater and consensus ratings. The process was repeated by independent raters to create a second set of individual and consensus ratings. Reliability of ratings of PEDro scale items was calculated using multi-rater kappas, and reliability of the total (summed) score was calculated using intraclass correlation coefficients (ICC [ 1, 1]). Results. The kappa value for each of the I I items ranged from .36 to .80 for individual assessors and from .50 to .79 for consensus ratings generated by groups of 2 or 3 raters. The ICC for the total score was .56 (95% confidence interval = .47-.65) for ratings by individuals, and the ICC for consensus ratings was .68 (95% confidence interval = .57-.76). Discussion and Conclusion. The reliability of ratings of PEDro scale items varied from "fair" to "substantial," and the reliability of the total PEDro score was "fair" to "good.".