Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system.
Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system.
复制标题
DOI:
10.1186/s12874-022-01649-y
复制
发表时间:
2022-06-08
影响因子:
4
通讯作者:
Muller, Ashley Elizabeth
中科院分区:
文献类型:
--
作者:
Jardim, Patricia Sofia Jacobsen;Rose, Christopher James;Ames, Heather Melanie;Echavez, Jose Francisco Meneses;Van de Velde, Stijn;Muller, Ashley Elizabeth
关键词:
Machine learning and automation are increasingly used to make the evidence synthesis process faster and more responsive to policymakers’ needs. In systematic reviews of randomized controlled trials (RCTs), risk of bias assessment is a resource-intensive task that typically requires two trained reviewers. One function of RobotReviewer, an off-the-shelf machine learning system, is an automated risk of bias assessment. We assessed the feasibility of adopting RobotReviewer within a national public health institute using a randomized, real-time, user-centered study. The study included 26 RCTs and six reviewers from two projects examining health and social interventions. We randomized these studies to one of two RobotReviewer platforms. We operationalized feasibility as accuracy, time use, and reviewer acceptability. We measured accuracy by the number of corrections made by human reviewers (either to automated assessments or another human reviewer’s assessments). We explored acceptability through group discussions and individual email responses after presenting the quantitative results. Reviewers were equally likely to accept judgment by RobotReviewer as each other’s judgement during the consensus process when measured dichotomously; risk ratio 1.02 (95% CI 0.92 to 1.13; p = 0.33). We were not able to compare time use. The acceptability of the program by researchers was mixed. Less experienced reviewers were generally more positive, and they saw more benefits and were able to use the tool more flexibly. Reviewers positioned human input and human-to-human interaction as superior to even a semi-automation of this process. Despite being presented with evidence of RobotReviewer’s equal performance to humans, participating reviewers were not interested in modifying standard procedures to include automation. If further studies confirm equal accuracy and reduced time compared to manual practices, we suggest that the benefits of RobotReviewer may support its future implementation as one of two assessors, despite reviewer ambivalence. Future research should study barriers to adopting automated tools and how highly educated and experienced researchers can adapt to a job market that is increasingly challenged by new technologies. The online version contains supplementary material available at 10.1186/s12874-022-01649-y.
登录
查看更多内容
DOI:
10.18653/v1/p17-4002
发表时间:
2017-07
期刊:
Proceedings of the conference. Association for Computational Linguistics. Meeting
影响因子:
--
作者:
Marshall IJ;Kuiper J;Banner E;Wallace BC
通讯作者:
Wallace BC
DOI:
10.1093/jamia/ocv044
发表时间:
2016-01-01
影响因子:
6.4
作者:
Marshall, Iain J.;Kuiper, Joel;Wallace, Byron C.
通讯作者:
Wallace, Byron C.
影响因子:
4
作者:
Ames, Heather;Glenton, Claire;Lewin, Simon
通讯作者:
Lewin, Simon
影响因子:
2.9
作者:
Borah R;Brown AW;Capers PL;Kaiser KA
通讯作者:
Kaiser KA
影响因子:
3.7
作者:
Armijo-Olivo S;Ospina M;da Costa BR;Egger M;Saltaji H;Fuentes J;Ha C;Cummings GG
通讯作者:
Cummings GG