Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews.

Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews.
复制标题

DOI:
10.1016/j.jclinepi.2020.11.003
复制
发表时间:
2021-05
影响因子:
7.2
通讯作者:
Marshall IJ
Marshall IJ
中科院分区:
医学2区
文献类型:
--
作者:
Thomas J;McDonald S;Noel-Storr A;Shemilt I;Elliott J;Mavergames C;Marshall IJ

文献摘要

参考文献

被引文献

相似文献

本研究开发、校准和评估了一种机器学习分类器,该分类器旨在减少科克伦系统评价中的研究识别工作量。开发了一种用于检索随机对照试验(RCT)的机器学习分类器(“科克伦RCT分类器”),该算法使用来自Embase的标题摘要记录数据集进行训练,由科克伦人群手动标记。然后,使用临床对冲团队手动标记的类似记录的进一步数据集校准分类器,目标是99%的召回率。最后,使用包含在科克伦综述中的RCT记录评估校准分类器的召回率,这些RCT具有足够长度的摘要以允许机器分类。科克伦RCT分类器使用280,620条记录(其中20,454条报告了RCT)进行训练。使用49,025个校准记录(其中1,587个报告了RCT)设置分类阈值,我们的bootstrap验证发现分类器在该数据集中的召回率为0.99(95%置信区间0.98-0.99),精度为0.08(95%置信区间0.06-0.12)。最终校准的随机对照试验分类器正确检索到了科克伦评论中包含的44,007项随机对照试验中的43,783项(99.5%),但错过了224项(0.5%)。较早的记录比最近发表的记录更容易被遗漏。科克伦随机对照试验分类器可以减少科克伦评价的人工研究识别工作量,缺失合格随机对照试验的风险非常低且可接受。该分类器现在是证据管道的一部分,证据管道是科克伦中部署的一个集成工作流程,有助于提高支持系统综述生成的研究识别流程的效率。系统性审查程序需要变得更有效率。机器学习已经足够成熟,可以在现实世界中使用。使用来自科克伦人群的数据构建机器学习分类器。它被校准以实现非常高的召回率。它现在已经上线并在科克伦审查生产系统中使用。
This study developed, calibrated, and evaluated a machine learning classifier designed to reduce study identification workload in Cochrane for producing systematic reviews. A machine learning classifier for retrieving randomized controlled trials (RCTs) was developed (the “Cochrane RCT Classifier”), with the algorithm trained using a data set of title–abstract records from Embase, manually labeled by the Cochrane Crowd. The classifier was then calibrated using a further data set of similar records manually labeled by the Clinical Hedges team, aiming for 99% recall. Finally, the recall of the calibrated classifier was evaluated using records of RCTs included in Cochrane Reviews that had abstracts of sufficient length to allow machine classification. The Cochrane RCT Classifier was trained using 280,620 records (20,454 of which reported RCTs). A classification threshold was set using 49,025 calibration records (1,587 of which reported RCTs), and our bootstrap validation found the classifier had recall of 0.99 (95% confidence interval 0.98–0.99) and precision of 0.08 (95% confidence interval 0.06–0.12) in this data set. The final, calibrated RCT classifier correctly retrieved 43,783 (99.5%) of 44,007 RCTs included in Cochrane Reviews but missed 224 (0.5%). Older records were more likely to be missed than those more recently published. The Cochrane RCT Classifier can reduce manual study identification workload for Cochrane Reviews, with a very low and acceptable risk of missing eligible RCTs. This classifier now forms part of the Evidence Pipeline, an integrated workflow deployed within Cochrane to help improve the efficiency of the study identification processes that support systematic review production. Systematic review processes need to become more efficient. Machine learning is sufficiently mature for real-world use. A machine learning classifier was built using data from Cochrane Crowd. It was calibrated to achieve very high recall. It is now live and in use in Cochrane review production systems.
DOI: 10.1111/j.1471-1842.2008.00827.x
发表时间: 2009-09-01
影响因子: 3.8
作者:
McKibbon, Kathleen Ann;Wilczynski, Nancy Lou;Haynes, Robert Brian
通讯作者: Haynes, Robert Brian
DOI: 10.1186/1472-6947-5-20
发表时间: 2005-06-21
影响因子: 3.5
作者:
Wilczynski, Nancy L;Morgan, Douglas;Haynes, R Brian
通讯作者: Haynes, R Brian
DOI: 10.1016/j.jclinepi.2020.08.008
发表时间: 2020-11-01
影响因子: 7.2
作者:
Noel-Storr, A. H.;Dooley, G.;Foxlee, R.
通讯作者: Foxlee, R.
DOI: 10.1371/journal.pmed.1000326
发表时间: 2010-09-21
期刊: PLoS medicine
影响因子: 15.8
作者:
Bastian H;Glasziou P;Chalmers I
通讯作者: Chalmers I
DOI: 10.1016/s0895-4356(01)00341-9
发表时间: 2001-08-01
影响因子: 7.2
作者:
Steyerberg, EW;Harrell, FE;Habbema, JDF
通讯作者: Habbema, JDF