Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach

Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach
复制标题

DOI:
10.1093/jamia/ocx053
复制
发表时间:
2017-11-01
影响因子:
6.4
通讯作者:
Thomas, James
Thomas, James
中科院分区:
管理学2区
文献类型:
--
作者:
Wallace, Byron C.;Noel-Storr, Anna;Thomas, James

文献摘要

被引文献

相似文献

识别所有已发表的随机对照试验(RCT)报告是一个重要的目标,但即使使用当前的机器学习(ML)方法,也需要大量的手动工作来区分RCT和非RCT。我们的目标是通过使用众包和机器学习的混合方法使这一过程更加高效。我们训练了一个分类器来区分描述随机对照试验的引文和不描述随机对照试验的引文。然后,我们采用了一种简单的策略,即自动排除被分类器认为不太可能是RCT的引文,否则将其提交给众包。(我们的估计显示95%-99%的回忆率),(我们观察到减少了大约60%-80%)。混合人群ML策略保证了生物医学策展/注释任务的进一步探索。
Identifying all published reports of randomized controlled trials (RCTs) is an important aim, but it requires extensive manual effort to separate RCTs from non-RCTs, even using current machine learning (ML) approaches. We aimed to make this process more efficient via a hybrid approach using both crowdsourcing and ML.We trained a classifier to discriminate between citations that describe RCTs and those that do not. We then adopted a simple strategy of automatically excluding citations deemed very unlikely to be RCTs by the classifier and deferring to crowdworkers otherwise.Combining ML and crowdsourcing provides a highly sensitive RCT identification strategy (our estimates suggest 95%-99% recall) with substantially less effort (we observed a reduction of around 60%-80%) than relying on manual screening alone.Hybrid crowd-ML strategies warrant further exploration for biomedical curation/annotation tasks.