Automated acquisition of disease-drug knowledge from biomedical and clinical documents: An initial study

Automated acquisition of disease-drug knowledge from biomedical and clinical documents: An initial study
复制标题

DOI:
10.1197/jamia.m2401
复制
发表时间:
2008-01-01
影响因子:
6.4
通讯作者:
Friedman, Carol
Friedman, Carol
中科院分区:
管理学2区
文献类型:
--
作者:
Chen, Elizabeth S.;Hripcsak, George;Friedman, Carol

文献摘要

被引文献

相似文献

目的:探索利用文本挖掘和统计技术自动获取生物医学和临床文献中的知识,以识别疾病-药物关联。设计:从患者病历中挖掘生物医学文献和临床叙述,以收集有关疾病-药物关联的知识。两个NLP系统BioMedLEE和MedLEE分别应用于Medline文章和出院摘要。除了Medline文章的网状注释外,还使用NLP系统识别疾病和药物实体。以8种疾病为研究对象,应用共现统计法计算和评价各疾病与相关药物之间的关联强度。结果:生成了疾病-药物对的排序列表,并计算了这些疾病-药物对之间较强关联的分界值,以便进一步分析。在疾病药物知识的文本来源(即生物医学文献和患者记录)和注释(即MESH和NLP提取的UMLS概念)方面存在差异和相似之处。结论:本文提出了一种获取疾病特定知识的方法,并对该方法的可行性进行了研究。该方法基于对生物医学和临床文献应用自然语言处理和统计技术的组合。该方法能够根据患者记录提取临床医生为特定疾病患者使用的药物的知识,同时也获得了经常参与这些疾病对照试验的药物的知识。在比较疾病-药物关联时,我们发现结果是适当的:两个文本来源包含一致和互补的知识,医学专家对排名前五的疾病-药物关联进行手动审查,支持它们在疾病中的正确性。
Objective: Explore the automated acquisition of knowledge in biomedical and clinical documents using text mining and statistical techniques to identify disease-drug associations.Design: Biomedical literature and clinical narratives from the patient record were mined to gather knowledge about disease-drug associations. Two NLP systems, BioMedLEE and MedLEE, were applied to Medline articles and discharge summaries, respectively. Disease and drug entities were identified using the NLP systems in addition to MeSH annotations for the Medline articles. Focusing on eight diseases, co-occurrence statistics were applied to compute and evaluate the strength of association between each disease and relevant drugs.Results: Ranked lists of disease-drug pairs were generated and cutoffs calculated for identifying stronger associations among these pairs for further analysis. Differences and similarities between the text sources (i.e., biomedical literature and patient record) and annotations (i.e., MeSH and NLP-extracted UMLS concepts) with regards to disease-drug knowledge were observed.Conclusion: This paper presents a method for acquiring disease-specific knowledge and a feasibility study of the method. The method is based on applying a combination of NLP and statistical techniques to both biomedical and clinical documents. The approach enabled extraction of knowledge about the drugs clinicians are using for patients with specific diseases based on the patient record, while it is also acquired knowledge of drugs frequently involved in controlled trials for those same diseases. In comparing the disease-drug associations, we found the results to be appropriate: the two text sources contained consistent as well as complementary knowledge, and manual review of the top five disease-drug associations by a medical expert supported their correctness across the diseases.