Towards Automatic Recognition of Scientifically Rigorous Clinical Research Evidence

Towards Automatic Recognition of Scientifically Rigorous Clinical Research Evidence
复制标题

DOI:
10.1197/jamia.m2996
复制
发表时间:
2009-01-01
影响因子:
6.4
通讯作者:
Haynes, R. Brian
Haynes, R. Brian
中科院分区:
管理学2区
文献类型:
--
作者:
Kilicoglu, Halil;Demner-Fushman, Dina;Haynes, R. Brian

文献摘要

被引文献

相似文献

由于文献检索方法的进步,越来越多的与主题相关的生物医学出版物随时可用,这对从事循证医学的临床医生构成了挑战。获取和批判性地评估现有证据变得越来越耗时。如果有方法可以自动识别在特定临床情况下立即适用的严格研究,这个问题可以在一定程度上得到解决。我们将从检索到的主题相关文章中识别包含有用临床建议的研究作为一个二进制分类问题。PubMed临床查询过滤器开发中使用的黄金标准构成了我们方法的基础。我们使用在高级语义特征上训练的有监督的机器学习技术(朴素贝叶斯、支持向量机和Boosting)来识别科学严谨的研究。我们使用集成学习方法(堆叠)将这些方法结合起来。学习方法的性能是使用准确率、召回率和F分数,以及接收者操作特征(ROC)曲线下的面积(AUC)来评估的。使用10,000条人工标注的MEDLINE引文的训练集和另外2,000条引文的测试集,我们在识别严格的、临床相关的研究方面获得了73.7%的准确率和61.5%的召回率,在堆叠超过5个特征-分类器组合的情况下,在识别具有治疗重点的严谨研究时,使用堆叠在单词+元数据特征向量上的准确率和召回率分别达到82.5%和84.3%。我们的结果表明,高质量的金标准和先进的分类方法可以帮助临床医生从医学文献中获得最好的证据。
The growing numbers of topically relevant biomedical publications readily available due to advances in document retrieval methods pose a challenge to clinicians practicing evidence-based medicine. It is increasingly time consuming to acquire and critically appraise the available evidence. This problem could be addressed in part if methods were available to automatically recognize rigorous studies immediately applicable in a specific clinical situation. We approach the problem of recognizing studies containing useable clinical advice from retrieved topically relevant articles as a binary classification problem. The gold standard used in the development of PubMed clinical query filters forms the basis of our approach. We identify scientifically rigorous studies using supervised machine learning techniques (Naive Bayes, support vector machine (SVM), and boosting) trained on high-level semantic features. We combine these methods using an ensemble learning method (stacking). The performance of learning methods is evaluated using precision, recall and F, score, in addition to area under the receiver operating characteristic (ROC) curve (AUC). Using a training set of 10,000 manually annotated MEDLINE citations, and a test set of an additional 2,000 citations, we achieve 73.7% precision and 61.5% recall in identifying rigorous, clinically relevant studies, with stacking over five feature-classifier combinations and 82.5% precision and 84.3% recall in recognizing rigorous studies with treatment focus using stacking over word + metadata feature vector. Our results demonstrate that a high quality gold standard and advanced classification methods can help clinicians acquire best evidence from the medical literature.