A statistical approach to scanning the biomedical literature for pharmacogenetics knowledge

A statistical approach to scanning the biomedical literature for pharmacogenetics knowledge
复制标题

DOI:
10.1197/jamia.m1640
复制
发表时间:
2005-03-01
影响因子:
6.4
通讯作者:
Altman, RB
Altman, RB
中科院分区:
管理学2区
文献类型:
--
作者:
Rubin, DL;Thorn, CF;Altman, RB

文献摘要

被引文献

相似文献

目的:生物医学数据库总结了当前的科学知识,但它们通常需要多年艰苦的管理工作来建立,重点是识别大量生物医学文献中的相关文献和数据。人工提取嵌入在大量文献中的有用信息是很困难的,自动智能文本分析工具在协助这些管理活动中变得越来越重要。作者的目标是开发一种自动化的方法来识别Medline引文中包含与基因-药物关系相关的药物遗传学数据的文章。设计:作者建立并评估了几个候选统计模型,这些模型根据词汇使用和这些文章中使用的医学主题词(MeSH)的概况来表征药物遗传学文章。使用表现最好的模型扫描整个Medline文章数据库(1100万篇文章)来识别候选药物遗传学文章。结果:一名药理学家审查了从扫描Medline中识别的文章样本,以评估该方法的精度。作者的方法在文献中确定了4892篇药物遗传学文章,准确率为92%。与人工积累所需的时间相比,他们的自动化方法只花了一小部分时间来获取这些文章。作者建立了一个Web资源(http://pharmdemo.stanford)。edu/pharmclb/main。间谍)提供访问他们的结果。结论:采用统计分类方法对药物遗传学文献进行筛选,准确度较高。这些方法可以帮助策展人获取建立生物医学数据库的相关文献。
Objective: Biomedical databases summarize current scientific knowledge, but they generally require years of laborious curation effort to build, focusing on identifying pertinent literature and data in the voluminous biomedical literature. It is difficult to manually extract useful information embedded in the large volumes of literature, and automated intelligent text analysis tools are becoming increasingly essential to assist in these curation activities. The goal of the authors was to develop an automated method to identify articles in Medline citations that contain pharmacogenetics data pertaining to gene-drug relationships.Design: The authors built and evaluated several candidate statistical models that characterize pharmacogenetics articles in terms of word usage and the profile of Medical Subject Headings (MeSH) used in those articles. The best-performing model was used to scan the entire Medline article database (11 million articles) to identify candidate pharmacogenetics articles.Results: A sampling of the articles identified from scanning Medline was reviewed by a pharmacologist to assess the precision of the method. The authors' approach identified 4,892 pharmacogenetics articles in the literature with 92% precision. Their automated method took a fraction of the time to acquire these articles compared with the time expected to be taken to accumulate them manually. The authors have built a Web resource (http://pharmdemo.stanford. edu/pharmclb/main.spy) to provide access to their results.Conclusion: A statistical classification approach can screen the primary literature to pharmacogenetics articles with high precision. Such methods may assist curators in acquiring pertinent literature in building biomedical databases.