Text categorization models for high-quality article retrieval in internal medicine

Text categorization models for high-quality article retrieval in internal medicine
复制标题

DOI:
10.1197/jamia.m1641
复制
发表时间:
2005-03-01
影响因子:
6.4
通讯作者:
Aliferis, CF
Aliferis, CF
中科院分区:
管理学2区
文献类型:
--
作者:
Aphinyanaphongs, Y;Tsamardinos, I;Aliferis, CF

文献摘要

被引文献

相似文献

目的:由于医学出版物的指数增长,找到适用于患者问题的最佳科学证据变得非常困难。本研究的目的是应用机器学习技术自动识别内科一段时间内的高质量、内容特定的文章,并将其性能与Haynes等人先前基于布尔的PubMed临床查询过滤器进行比较。设计:ACP杂志俱乐部对内科文章的选择标准是在病因学、预后、诊断和治疗。应用Naive Bayes(一种专门的AdaBoost算法)以及线性和多项式支持向量机来识别这些文章。测量:机器学习模型在每个类别中相互比较,并与临床查询过滤器进行比较,使用接收器工作特征曲线下的面积,11点平均召回精度和灵敏度/特异性匹配方法。结果:在大多数类别中,数据诱导模型具有比临床查询过滤器更好或相当的灵敏度、特异性和精确度。多项式支持向量机模型在所有学习方法中表现最好,通过受试者工作曲线下面积和11点平均召回率对文章进行排名。这项研究表明,使用机器学习方法,可以自动构建检索高质量,内容特定的文章,使用ACP杂志俱乐部的纳入或引用作为金标准,在给定的时间段内,在内科学中表现优于1994年PubMed临床查询过滤器。
Objective: Finding the best scientific evidence that applies to a patient problem is becoming exceedingly difficult due to the exponential growth of medical publications. The objective of this study was to apply machine learning techniques to automatically identify high-quality, content-specific articles for one time period in internal medicine and compare their performance with previous Boolean-based PubMed clinical query filters of Haynes et al.Design: The selection criteria of the ACP journal Club for articles in internal medicine were the basis for identifying high-quality articles in the areas of etiology, prognosis, diagnosis, and treatment. Naive Bayes, a specialized AdaBoost algorithm, and linear and polynomial support vector machines were applied to identify these articles.Measurements: The machine learning models were compared in each category with each other and with the clinical query filters using area under the receiver operating characteristic curves, 11-point average recall precision, and a sensitivity/specificity match method.Results: In most categories, the data-induced models have better or comparable sensitivity, specificity, and precision than the clinical query filters. The polynomial support vector machine models perform the best among all learning methods in ranking the articles as evaluated by area under the receiver operating curve and 11-point average recall precision.Conclusion: This research shows that, using machine learning methods, it is possible to automatically build models for retrieving high-quality, content-specific articles using inclusion or citation by the ACP journal Club as a gold standard in a given time period in internal medicine that perform better than the 1994 PubMed clinical query filters.