Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach
Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach
复制标题
DOI:
10.1093/jamia/ocx053
复制
发表时间:
2017-11-01
影响因子:
6.4
通讯作者:
Thomas, James
中科院分区:
文献类型:
--
作者:
Wallace, Byron C.;Noel-Storr, Anna;Thomas, James
Identifying all published reports of randomized controlled trials (RCTs) is an important aim, but it requires extensive manual effort to separate RCTs from non-RCTs, even using current machine learning (ML) approaches. We aimed to make this process more efficient via a hybrid approach using both crowdsourcing and ML.We trained a classifier to discriminate between citations that describe RCTs and those that do not. We then adopted a simple strategy of automatically excluding citations deemed very unlikely to be RCTs by the classifier and deferring to crowdworkers otherwise.Combining ML and crowdsourcing provides a highly sensitive RCT identification strategy (our estimates suggest 95%-99% recall) with substantially less effort (we observed a reduction of around 60%-80%) than relying on manual screening alone.Hybrid crowd-ML strategies warrant further exploration for biomedical curation/annotation tasks.