Equivalence of electronic and paper administration of patient-reported outcome measures: a systematic review and meta-analysis of studies conducted between 2007 and 2013

Equivalence of electronic and paper administration of patient-reported outcome measures: a systematic review and meta-analysis of studies conducted between 2007 and 2013
复制标题

DOI:
10.1186/s12955-015-0362-x
复制
发表时间:
2015-10-07
影响因子:
3.6
通讯作者:
Wild, Diane J.
Wild, Diane J.
中科院分区:
医学3区
文献类型:
--
作者:
Muehlhausen, Willie;Doll, Helen;Wild, Diane J.

文献摘要

被引文献

相似文献

目的:对 Gwaltney 等人 2008 年综述中纳入的研究中患者报告结果测量 (PROM) 的电子和纸质管理之间的等效性进行系统回顾和荟萃分析。方法:对 2007 年至 2013 年间进行的 PROM 等效性研究进行系统文献回顾,确定了 1,997 条记录,其中 72 项研究符合预先定义的纳入/排除标准。根据相关系数(ICC、Spearman 和 Pearson 相关性、Kappa 统计)和平均差异(通过标准差、SD 和响应量表范围标准化)提取每项研究的 PRO 数据。估计相关性和平均差的汇总估计值。检查了给药方式、发表年份、研究设计、给药之间的时间间隔、参与者的平均年龄和出版物类型的修正效应。结果:提取了 435 个个体相关性,这些相关性变化很大 (I2 = 93.8),但总体显示出良好的等效性,ICC 范围为 0.65 至 0.99,汇总相关系数为 0.88(95% CI 0.87 至 0.87)。 0.88)。 307 项研究的标准化均值差异较小且变量较少 (I2 = 33.5),汇总标准化均值差异为 0.037(95% CI 0.031 至 0.042)。 56 项研究(61 项估计值)的平均给药模式/平台特定相关性的汇总估计值为 0.88(95% CI 0.86 至 0.90),并且仍然存在很大差异(I2 = 92.1)。同样,来自 39 项研究(42 项估计值)的平均平台特定 ICC 的汇总估计值为 0.90(95% CI 0.88 至 0.92),I2 为 91.5。排除 20 项相关系数偏远(距平均值 >= 3SD)的研究后,I2 为 54.4,等效性仍然很高,总体汇总相关系数为 0.88(95% CI 0.87 至 0.88)。最近的研究(p < 0.001)、随机研究与非随机研究(p < 0.001)、间隔较短(< 1 天)的研究(p < 0.001)以及平均年龄 28 至 55 岁的受访者与年轻或年长受访者相比(p < 0.001)的一致性更高。就模式/平台而言,纸质与交互式语音应答系统 (IVRS) 的比较具有最低的汇总一致性,纸质与平板电脑/触摸屏的比较最高 (p < 0.001)。结论:本研究支持 Gwaltney 之前的荟萃分析的结论,即纸质 PROM 与电子设备管理的测量在数量上具有可比性。它还证实了 ISPOR 特别工作组的结论,即仅进行微小变化的迁移不需要进行定量等效研究。这一发现应该会让调查人员、监管机构和申办者在使用最佳实践进行迁移后在电子设备上进行调查问卷感到放心。尽管有数据表明适度变化的迁移会产生等效的仪器版本,因此不需要定量等效性研究,但需要进行额外的工作来确定这一点。此外,需要标准化迁移实践和报告实践(即包括测试仪器版本和屏幕截图的副本),以便将来可以就等效性测试提出明确的建议。提出有关继续进行等效性测试的必要性的问题。
Objective: To conduct a systematic review and meta-analysis of the equivalence between electronic and paper administration of patient reported outcome measures (PROMs) in studies conducted subsequent to those included in Gwaltney et al's 2008 review.Methods: A systematic literature review of PROM equivalence studies conducted between 2007 and 2013 identified 1,997 records from which 72 studies met pre-defined inclusion/exclusion criteria. PRO data from each study were extracted, in terms of both correlation coefficients (ICCs, Spearman and Pearson correlations, Kappa statistics) and mean differences (standardized by the standard deviation, SD, and the response scale range). Pooled estimates of correlation and mean difference were estimated. The modifying effects of mode of administration, year of publication, study design, time interval between administrations, mean age of participants and publication type were examined.Results: Four hundred thirty-five individual correlations were extracted, these correlations being highly variable (I2 = 93.8) but showing generally good equivalence, with ICCs ranging from 0.65 to 0.99 and the pooled correlation coefficient being 0.88 (95 % CI 0.87 to 0.88). Standardised mean differences for 307 studies were small and less variable (I2 = 33.5) with a pooled standardised mean difference of 0.037 (95 % CI 0.031 to 0.042). Average administration mode/platform-specific correlations from 56 studies (61 estimates) had a pooled estimate of 0.88 (95 % CI 0.86 to 0.90) and were still highly variable (I2 = 92.1). Similarly, average platform-specific ICCs from 39 studies (42 estimates) had a pooled estimate of 0.90 (95 % CI 0.88 to 0.92) with an I2 of 91.5. After excluding 20 studies with outlying correlation coefficients (>= 3SD from the mean), the I2 was 54.4, with the equivalence still high, the overall pooled correlation coefficient being 0.88 (95 % CI 0.87 to 0.88). Agreement was found to be greater in more recent studies (p < 0.001), in randomized studies compared with non-randomised studies (p < 0.001), in studies with a shorter interval (< 1 day) (p < 0.001), and in respondents of mean age 28 to 55 compared with those either younger or older (p < 0.001). In terms of mode/platform, paper vs Interactive Voice Response System (IVRS) comparisons had the lowest pooled agreement and paper vs tablet/touch screen the highest (p < 0.001).Conclusion: The present study supports the conclusion of Gwaltney's previous meta-analysis showing that PROMs administered on paper are quantitatively comparable with measures administered on an electronic device. It also confirms the ISPOR Taskforce's conclusion that quantitative equivalence studies are not required for migrations with minor change only. This finding should be reassuring to investigators, regulators and sponsors using questionnaires on electronic devicesafter migration using best practices. Although there is data indicating that migrations with moderate changes produce equivalent instrument versions, hence do not require quantitative equivalence studies, additional work is necessary to establish this. Furthermore, there is the need to standardize migration practices and reporting practices (i.e. include copies of tested instrument versions and screenshots) so that clear recommendations regarding equivalence testing can be made in the future. raising questions about the necessity of conducting equivalence testing moving forward.