Using citation data to improve retrieval from MEDLINE

Using citation data to improve retrieval from MEDLINE
复制标题

DOI:
10.1197/jamia.m1909
复制
发表时间:
2006-01-01
影响因子:
6.4
通讯作者:
Hersh, WR
Hersh, WR
中科院分区:
管理学2区
文献类型:
--
作者:
Bernstam, EV;Herskovic, JR;Hersh, WR

文献摘要

被引文献

相似文献

目标:确定为万维网开发的算法是否可以应用于生物医学文献,以便识别重要且相关的文章。设计和测量:直接比较八种算法:简单的 PubMed 查询、临床查询(敏感和特定版本)、向量余弦比较、引用计数、期刊影响因子、PageRank 和基于多项式支持向量的机器学习 机器。目的是对重要文章进行优先排序,重要文章被定义为包含在外科肿瘤学重要文献的现有参考书目中。结果:基于引文的算法在识别重要文章方面比基于非引文的算法更有效。最有效的策略是简单的引用计数和 PageRank,它平均在前 100 个结果中识别出超过 6 篇重要文章,而最佳基于非引用的算法为 0.85 (p < 0.001)。作者在 10、20、50、200、500 和 1,000 个结果中发现基于引用和基于非引用的算法之间存在类似差异 (p < 0.001)。引用滞后对 PageRank 性能的影响比简单的引用计数更大。然而,尽管存在引文滞后,基于引文的算法仍然比基于非引文的算法更有效。结论:在万维网上证明成功的算法可以应用于生物医学信息检索。基于引文的算法可以帮助识别大量相关结果中的重要文章。需要进一步的研究来确定基于引文的算法是否能够有效满足用户的实际信息需求。
Objective: To determine whether algorithms developed for the World Wide Web can be applied to the biomedical literature in order to identify articles that are important as well as relevant.Design and Measurements: A direct comparison of eight algorithms: simple PubMed queries, clinical queries (sensitive and specific versions), vector cosine comparison, citation count, journal impact factor, PageRank, and machine learning based on polynomial support vector machines. The objective was to prioritize important articles, defined as being included in a pre-existing bibliography of important literature in surgical oncology.Results: Citation-based algorithms were more effective than noncitation-based algorithms at identifying important articles. The most effective strategies were simple citation count and PageRank, which on average identified over six important articles in the first 100 results compared to 0.85 for the best noncitation-based algorithm (p < 0.001). The authors saw similar differences between citation-based and noncitation-based algorithms at 10, 20, 50, 200, 500, and 1,000 results (p < 0.001). Citation lag affects performance of PageRank more than simple citation count. However, in spite of citation lag, citation-based algorithms remain more effective than noncitation-based algorithms.Conclusion: Algorithms that have proved successful on the World Wide Web can be applied to biomedical information retrieval. Citation-based algorithms can help identify important articles within large sets of relevant results. Further studies are needed to determine whether citation-based algorithms can effectively meet actual user information needs.