Integrating existing natural language processing tools for medication extraction from discharge summaries

Integrating existing natural language processing tools for medication extraction from discharge summaries
复制标题

DOI:
10.1136/jamia.2010.003855
复制
发表时间:
2010-09-01
影响因子:
6.4
通讯作者:
Xu, Hua
Xu, Hua
中科院分区:
管理学2区
文献类型:
--
作者:
Doan, Son;Bastarache, Lisa;Xu, Hua

文献摘要

被引文献

相似文献

目的开发一个自动化系统,从出院小结中提取药物和相关信息,作为2009年i2 b2自然语言处理(NLP)挑战赛的一部分。这项任务需要准确识别的药物名称,剂量,模式,频率,持续时间和原因给药administration.Design我们开发了一个集成系统,使用现有的几个NLP组件开发在范德比尔特大学医学中心,其中包括MedEx(提取药物信息),Sec标签(一节识别系统的临床笔记),一个句子分割器,和拼写检查的药物名称。我们的目标是在对这个文档语料库进行最少或没有特定训练的情况下实现良好的性能;因此,评估这些NLP工具在其所在机构之外的可移植性。综合系统使用了由主办方标注的17个注释,并使用参赛团队标注的251个注释进行了评估。测量i2 b2挑战赛使用标准测量方法,包括精确度,召回率和F-测量,以评估参赛系统的性能。有两种方法可以确定提取的文本结果是否正确:精确匹配或不精确匹配。所有六种类型的药物相关的研究结果在251注释notes的整体性能被认为是在challenge.Results的主要指标,我们的系统实现了整体F-测量为0.821精确匹配(0.839精度; 0.803召回)和0.822不精确匹配(0.866精度; 0.782召回)。该系统排名第二的20个参与团队的整体性能在提取药物和相关information.Conclusions的结果表明,现有的MedEx系统,连同其他NLP组件,可以提取药物信息的临床文本从机构以外的网站的算法开发具有合理的性能。
Objective To develop an automated system to extract medications and related information from discharge summaries as part of the 2009 i2b2 natural language processing (NLP) challenge. This task required accurate recognition of medication name, dosage, mode, frequency, duration, and reason for drug administration.Design We developed an integrated system using several existing NLP components developed at Vanderbilt University Medical Center, which included MedEx (to extract medication information), Sec Tag (a section identification system for clinical notes), a sentence splitter, and a spell checker for drug names. Our goal was to achieve good performance with minimal to no specific training for this document corpus; thus, evaluating the portability of those NLP tools beyond their home institution. The integrated system was developed using 17 notes that were annotated by the organizers and evaluated using 251 notes that were annotated by participating teams.Measurements The i2b2 challenge used standard measures, including precision, recall, and F-measure, to evaluate the performance of participating systems. There were two ways to determine whether an extracted textual finding is correct or not: exact matching or inexact matching. The overall performance for all six types of medication-related findings across 251 annotated notes was considered as the primary metric in the challenge.Results Our system achieved an overall F-measure of 0.821 for exact matching (0.839 precision; 0.803 recall) and 0.822 for inexact matching (0.866 precision; 0.782 recall). The system ranked second out of 20 participating teams on overall performance at extracting medications and related information.Conclusions The results show that the existing MedEx system, together with other NLP components, can extract medication information in clinical text from institutions other than the site of algorithm development with reasonable performance.