Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system.

Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system.
复制标题

DOI:
10.1186/s12874-022-01649-y
复制
发表时间:
2022-06-08
影响因子:
4
通讯作者:
Muller, Ashley Elizabeth
Muller, Ashley Elizabeth
中科院分区:
医学3区
文献类型:
--
作者:
Jardim, Patricia Sofia Jacobsen;Rose, Christopher James;Ames, Heather Melanie;Echavez, Jose Francisco Meneses;Van de Velde, Stijn;Muller, Ashley Elizabeth

文献摘要

参考文献

被引文献

相似文献

机器学习和自动化越来越多地被用于使证据合成过程更快,更好地响应政策制定者的需求。在随机对照试验(RCT)的系统评价中,偏差风险评估是一项资源密集型任务,通常需要两名训练有素的评价者。RobotReviewer是一个现成的机器学习系统,它的一个功能是自动进行偏见风险评估。我们使用一项随机的、实时的、以用户为中心的研究,评估了在国家公共卫生机构内采用RobotReviewer的可行性。这项研究包括26名随机对照试验和来自两个检查卫生和社会干预措施的项目的6名评价者。我们将这些研究随机分配到两个RobotReviewer平台中的一个。我们将可行性视为精确度、时间使用和评审者的可接受性。我们通过人工审核者所做的更正次数来衡量准确性(无论是对自动评估还是对另一位人工审查者的评估)。在展示量化结果后,我们通过小组讨论和个人电子邮件回复来探索可接受性。评价者同样可能接受机器人评审员的判断,将其作为共识过程中彼此的判断;风险比为1.02(95%可信区间0.92比1.13;p = 0.33)。我们无法比较时间使用情况。研究人员对该项目的接受程度参差不齐。经验较少的评审员通常更积极,他们看到了更多的好处,能够更灵活地使用该工具。评价者认为,人工输入和人与人之间的互动甚至优于这一过程的半自动。尽管有证据表明RobotReviewer的表现与人类一样,但参与评审者对修改标准程序以包括自动化不感兴趣。如果进一步的研究证实了与人工实践相同的准确性和减少的时间,我们建议RobotReviewer的好处可能支持它作为两个评价者之一的未来实施,尽管评价者的矛盾心理。未来的研究应该研究采用自动化工具的障碍,以及受过高等教育和经验丰富的研究人员如何适应日益受到新技术挑战的就业市场。网上版载有补充材料,可在10.1186/s12874-022-01649-y查阅。
Machine learning and automation are increasingly used to make the evidence synthesis process faster and more responsive to policymakers’ needs. In systematic reviews of randomized controlled trials (RCTs), risk of bias assessment is a resource-intensive task that typically requires two trained reviewers. One function of RobotReviewer, an off-the-shelf machine learning system, is an automated risk of bias assessment. We assessed the feasibility of adopting RobotReviewer within a national public health institute using a randomized, real-time, user-centered study. The study included 26 RCTs and six reviewers from two projects examining health and social interventions. We randomized these studies to one of two RobotReviewer platforms. We operationalized feasibility as accuracy, time use, and reviewer acceptability. We measured accuracy by the number of corrections made by human reviewers (either to automated assessments or another human reviewer’s assessments). We explored acceptability through group discussions and individual email responses after presenting the quantitative results. Reviewers were equally likely to accept judgment by RobotReviewer as each other’s judgement during the consensus process when measured dichotomously; risk ratio 1.02 (95% CI 0.92 to 1.13; p = 0.33). We were not able to compare time use. The acceptability of the program by researchers was mixed. Less experienced reviewers were generally more positive, and they saw more benefits and were able to use the tool more flexibly. Reviewers positioned human input and human-to-human interaction as superior to even a semi-automation of this process. Despite being presented with evidence of RobotReviewer’s equal performance to humans, participating reviewers were not interested in modifying standard procedures to include automation. If further studies confirm equal accuracy and reduced time compared to manual practices, we suggest that the benefits of RobotReviewer may support its future implementation as one of two assessors, despite reviewer ambivalence. Future research should study barriers to adopting automated tools and how highly educated and experienced researchers can adapt to a job market that is increasingly challenged by new technologies. The online version contains supplementary material available at 10.1186/s12874-022-01649-y.
DOI: 10.18653/v1/p17-4002
发表时间: 2017-07
期刊: Proceedings of the conference. Association for Computational Linguistics. Meeting
影响因子: --
作者:
Marshall IJ;Kuiper J;Banner E;Wallace BC
通讯作者: Wallace BC
DOI: 10.1093/jamia/ocv044
发表时间: 2016-01-01
影响因子: 6.4
作者:
Marshall, Iain J.;Kuiper, Joel;Wallace, Byron C.
通讯作者: Wallace, Byron C.
DOI: 10.1186/s12874-019-0665-4
发表时间: 2019-01-31
影响因子: 4
作者:
Ames, Heather;Glenton, Claire;Lewin, Simon
通讯作者: Lewin, Simon
DOI: 10.1136/bmjopen-2016-012545
发表时间: 2017-02-27
期刊: BMJ open
影响因子: 2.9
作者:
Borah R;Brown AW;Capers PL;Kaiser KA
通讯作者: Kaiser KA
DOI: 10.1371/journal.pone.0096920
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Armijo-Olivo S;Ospina M;da Costa BR;Egger M;Saltaji H;Fuentes J;Ha C;Cummings GG
通讯作者: Cummings GG