Evaluation of an automatic article selection method for timelier updates of the Comet Core Outcome Set database

Evaluation of an automatic article selection method for timelier updates of the Comet Core Outcome Set database
复制标题

DOI:
10.1093/incorrectjnlcode/baz109
复制
发表时间:
2019-11-07
影响因子:
5.8
通讯作者:
Williamson, Paula R.
Williamson, Paula R.
中科院分区:
生物学4区
文献类型:
--
作者:
Norman, Christopher R.;Gargon, Elizabeth;Williamson, Paula R.

文献摘要

被引文献

相似文献

科学文献的精选数据库在帮助研究人员找到相关文献方面发挥着重要作用,但填充此类数据库是一个劳动密集型和耗时的过程。其中一个数据库是可免费访问的Comet Core Outcome Set数据库,该数据库最初是在每年更新的系统综述中使用手动筛选进行填充的。为了减少工作量并促进更及时的更新,我们正在评估机器学习方法,以减少筛选所需的参考文献数量。在这项研究中,我们评估了一种基于逻辑回归的机器学习方法来自动对候选文章进行排名。来自原始系统性综述及其四次首次综述更新的数据用于训练模型并评估性能。我们估计,使用自动筛选将产生至少75%的工作量减少,同时保持遗漏参考文献的数量在2%左右。我们认为这是一个可接受的权衡,为这个系统的审查,该方法目前正在使用的彗星数据库更新的下一轮。
Curated databases of scientific literature play an important role in helping researchers find relevant literature, but populating such databases is a labour intensive and time-consuming process. One such database is the freely accessible Comet Core Outcome Set database, which was originally populated using manual screening in an annually updated systematic review. In order to reduce the workload and facilitate more timely updates we are evaluating machine learning methods to reduce the number of references needed to screen. In this study we have evaluated a machine learning approach based on logistic regression to automatically rank the candidate articles. Data from the original systematic review and its four first review updates were used to train the model and evaluate performance. We estimated that using automatic screening would yield a workload reduction of at least 75% while keeping the number of missed references around 2%. We judged this to be an acceptable trade-off for this systematic review, and the method is now being used for the next round of the Comet database update.