Automated confidence ranked classification of randomized controlled trial articles: an aid to evidence-based medicine.

Automated confidence ranked classification of randomized controlled trial articles: an aid to evidence-based medicine.
复制标题

DOI:
10.1093/jamia/ocu025
复制
发表时间:
2015-05
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Yu PS
Yu PS
中科院分区:
其他
文献类型:
--
作者:
Cohen AM;Smalheiser NR;McDonagh MS;Yu C;Adams CE;Davis JM;Yu PS

文献摘要

参考文献

被引文献

相似文献

目的:对于许多文献综述任务,包括系统综述(SR)和循证医学的其他方面,了解一篇文章是否描述了随机对照试验(RCT)非常重要。目前的手动注释对于SR过程来说不够完整或灵活。在这项工作中,建立了高度准确的机器学习预测模型,包括对一篇文章是否为RCT的置信度预测。材料与方法:LibSVM分类器与MEDLINE的大型人类相关子集上的潜在特征集的前向选择一起使用,以创建仅需要每篇文章的引文、摘要和MeSH术语的分类模型。结果如下:该模型实现了0.973的受试者工作特征曲线下的面积和2011年数据的均方误差为0.013。在手动审查的一组供试品上确认了准确的置信度估计值。还创建了不需要MeSH术语的第二个模型,并且表现得几乎一样好。讨论:两种模型均准确排名和预测文章RCT置信度。使用该模型和手动审查的样本,估计可以在MEDLINE中识别约8000(3%)个额外的RCT,并且可能无法识别Medline中标记为RCT的文章的5%。结论:使用连续评估的RCT置信度重新标记人类相关研究可能比简单的是/否预测更有助于文章排名和审查。自动化RCT标记工具应该在编写SR的过程中节省大量的时间和精力,并且是我们正在构建的用于简化SR工作流程的多步骤文本挖掘管道的关键组成部分。此外,该模型可能有助于识别MEDLINE出版物类型中的错误。这里描述的RCT置信度预测已经作为Web服务提供给用户,用户查询表单前端位于:http://arrowsmith.psych.uic.edu/cgi-bin/arrowsmith_uic/RCT_Tagger.cgi。
Objective: For many literature review tasks, including systematic review (SR) and other aspects of evidence-based medicine, it is important to know whether an article describes a randomized controlled trial (RCT). Current manual annotation is not complete or flexible enough for the SR process. In this work, highly accurate machine learning predictive models were built that include confidence predictions of whether an article is an RCT. Materials and Methods: The LibSVM classifier was used with forward selection of potential feature sets on a large human-related subset of MEDLINE to create a classification model requiring only the citation, abstract, and MeSH terms for each article. Results: The model achieved an area under the receiver operating characteristic curve of 0.973 and mean squared error of 0.013 on the held out year 2011 data. Accurate confidence estimates were confirmed on a manually reviewed set of test articles. A second model not requiring MeSH terms was also created, and performs almost as well. Discussion: Both models accurately rank and predict article RCT confidence. Using the model and the manually reviewed samples, it is estimated that about 8000 (3%) additional RCTs can be identified in MEDLINE, and that 5% of articles tagged as RCTs in Medline may not be identified. Conclusion: Retagging human-related studies with a continuously valued RCT confidence is potentially more useful for article ranking and review than a simple yes/no prediction. The automated RCT tagging tool should offer significant savings of time and effort during the process of writing SRs, and is a key component of a multistep text mining pipeline that we are building to streamline SR workflow. In addition, the model may be useful for identifying errors in MEDLINE publication types. The RCT confidence predictions described here have been made available to users as a web service with a user query form front end at: http://arrowsmith.psych.uic.edu/cgi-bin/arrowsmith_uic/RCT_Tagger.cgi.
DOI: 10.1186/1472-6947-12-33
发表时间: 2012-04-19
影响因子: 3.5
作者:
Cohen AM;Ambert K;McDonagh M
通讯作者: McDonagh M
DOI: 10.1186/1472-6947-9-10
发表时间: 2009-01-10
影响因子: 3.5
作者:
Chung, Grace Y.
通讯作者: Chung, Grace Y.
DOI: 10.1197/jamia.m1641
发表时间: 2005-03-01
影响因子: 6.4
作者:
Aphinyanaphongs, Y;Tsamardinos, I;Aliferis, CF
通讯作者: Aliferis, CF
DOI: 10.1177/1049732312452938
发表时间: 2012-10-01
影响因子: 3.2
作者:
Cooke, Alison;Smith, Debbie;Booth, Andrew
通讯作者: Booth, Andrew
DOI: 10.1197/jamia.m2996
发表时间: 2009-01-01
影响因子: 6.4
作者:
Kilicoglu, Halil;Demner-Fushman, Dina;Haynes, R. Brian
通讯作者: Haynes, R. Brian