RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials

RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials
复制标题

DOI:
10.1093/jamia/ocv044
复制
发表时间:
2016-01-01
影响因子:
6.4
通讯作者:
Wallace, Byron C.
Wallace, Byron C.
中科院分区:
管理学2区
文献类型:
--
作者:
Marshall, Iain J.;Kuiper, Joel;Wallace, Byron C.

文献摘要

被引文献

相似文献

目的开发和评价一种自动评估临床试验偏倚的机器学习系统RobotReviewer。从一个(PDF格式)的试验报告,系统应确定风险的偏倚定义的领域的科克伦风险偏倚(RoB)tool,并提取支持这些judgment.Methods文本,我们算法注释12,808试验PDF使用的数据从科克伦数据库的系统评价(CDSR)。试验被标记为每个领域的偏倚风险低或高/不清楚,句子被标记为是否提供信息。该数据集用于训练多任务ML模型。我们通过比较CDSR中两个或多个独立RoB评估的试验,估计了ML判断与人类的准确性。20名经验丰富的盲法评审员对支持性文本的相关性进行了评级,并将ML输出与等效输出进行了比较结果通过检索每个文档的前3个候选句子,(top3召回),最佳ML文本被评为比CDSR文本更相关,但不显著(60.4%的ML文本被评为“高度相关”,56.5%的文本来自评论;差异为+3.9%,[-3.2%至+10.9%])。模型RoB判断的准确性低于发表的评论,尽管差异是
Objective To develop and evaluate RobotReviewer, a machine learning (ML) system that automatically assesses bias in clinical trials. From a (PDF-formatted) trial report, the system should determine risks of bias for the domains defined by the Cochrane Risk of Bias (RoB) tool, and extract supporting text for these judgments.Methods We algorithmically annotated 12,808 trial PDFs using data from the Cochrane Database of Systematic Reviews (CDSR). Trials were labeled as being at low or high/unclear risk of bias for each domain, and sentences were labeled as being informative or not. This dataset was used to train a multi-task ML model. We estimated the accuracy of ML judgments versus humans by comparing trials with two or more independent RoB assessments in the CDSR. Twenty blinded experienced reviewers rated the relevance of supporting text, comparing ML output with equivalent (human-extracted) text from the CDSR.Results By retrieving the top 3 candidate sentences per document (top3 recall), the best ML text was rated more relevant than text from the CDSR, but not significantly (60.4% ML text rated 'highly relevant' v 56.5% of text from reviews; difference +3.9%, [-3.2% to +10.9%]). Model RoB judgments were less accurate than those from published reviews, though the difference was