Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications

Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications
复制标题

DOI:
10.1136/jamia.2009.001560
复制
发表时间:
2010-09-01
影响因子:
6.4
通讯作者:
Chute, Christopher G.
Chute, Christopher G.
中科院分区:
管理学2区
文献类型:
--
作者:
Savova, Guergana K.;Masanz, James J.;Chute, Christopher G.

文献摘要

被引文献

相似文献

我们旨在构建和评估一种开源自然语言处理系统,以从电子病历临床自由文本中提取信息。我们在http://www.ohnlp.org上描述和评估我们的系统,临床文本分析和知识提取系统(CTAKE)。 CTAKE基于现有的开源技术,非结构化信息管理体系结构框架和OpenNLP自然语言处理工具包。它的组件,专门针对临床领域培训,创建了丰富的语言和语义注释。单个组件的性能:句子边界检测器精度= 0.949;令牌的精度= 0.949;语音标记的一部分精度= 0.936;浅解析器f-评分= 0.924;命名实体识别器和系统级评估f-Score = 0.715,重叠跨度为0.715,概念映射,否定和状态属性的准确性为0.957、0.943、0.859和0.580和0.580、0.939,以及0.943、0.859和0.939,以及0.943、0.859和0.939,以及分别为0.839。讨论了五个应用程序的总体绩效。 Ctakes注释是临床自由文本高级语义处理方法和模块的基础。
We aim to build and evaluate an open-source natural language processing system for information extraction from electronic medical record clinical free-text. We describe and evaluate our system, the clinical Text Analysis and Knowledge Extraction System (cTAKES), released open-source at http://www.ohnlp.org. The cTAKES builds on existing open-source technologies the Unstructured Information Management Architecture framework and OpenNLP natural language processing toolkit. Its components, specifically trained for the clinical domain, create rich linguistic and semantic annotations. Performance of individual components: sentence boundary detector accuracy=0.949; tokenizer accuracy=0.949; part-of-speech tagger accuracy=0.936; shallow parser F-score=0.924; named entity recognizer and system-level evaluation F-score=0.715 for exact and 0.824 for overlapping spans, and accuracy for concept mapping, negation, and status attributes for exact and overlapping spans of 0.957, 0.943, 0.859, and 0.580, 0.939, and 0.839, respectively. Overall performance is discussed against five applications. The cTAKES annotations are the foundation for methods and modules for higher-level semantic processing of clinical free-text.