Towards comprehensive syntactic and semantic annotations of the clinical narrative.

Towards comprehensive syntactic and semantic annotations of the clinical narrative.
复制标题

DOI:
10.1136/amiajnl-2012-001317
复制
发表时间:
2013-09
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Savova GK
Savova GK
中科院分区:
其他
文献类型:
--
作者:
Albright D;Lanfranchi A;Fredriksen A;Styler WF 4th;Warner C;Hwang JD;Choi JD;Dligach D;Nielsen RD;Martin J;Ward W;Palmer M;Savova GK

文献摘要

参考文献

被引文献

相似文献

创建具有句法和语义标签层的注释临床叙述,以促进临床自然语言处理(NLP)的进步。开发NLP算法和开源组件。手动注释的临床叙事语料库的127 606令牌以下的树库模式的句法信息,PropBank模式的谓词参数结构,和统一的医学语言系统(UMLS)模式的语义信息。 开发了NLP组件。最终语料库由13091个句子组成,包含1772个不同的谓词词元。 在新创建的766个PropBank框架中,有74个是动词。有28539命名实体(NE)注释分布在15个UMLS语义组,一个UMLS语义类型,和人的语义类别。 最常见的注释属于UMLS语义组程序(15.71%),疾病(14.74%),概念和想法(15.10%),解剖(12.80%),化学品和药物(7.49%),以及UMLS语义类型的体征或症状(12.46%)。注释器间一致性结果:Treebank(0.926),PropBank(0.891-0.931),NE(0.697-0.750)。词性标注器、选区解析器、依存关系解析器和语义角色标注器都是从语料库构建的,并开源发布。该项目发现的一个重要限制是NLP社区需要开发一个广泛认可的模式来注释临床概念及其关系。该项目迈出了基础性的一步,使临床NLP领域与一般领域的NLP相媲美。语料库创建和NLP组件为研究和应用程序开发提供了以前不可能的资源。
To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891–0.931), NE (0.697–0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible.
DOI: 10.1197/jamia.m1733
发表时间: 2005-05-01
影响因子: 6.4
作者:
Hripcsak, G;Rothschild, AS
通讯作者: Rothschild, AS
DOI: 10.1016/j.jbi.2008.12.013
发表时间: 2009-10-01
影响因子: 4.5
作者:
Roberts, Angus;Gaizauskas, Robert;Setzer, Andrea
通讯作者: Setzer, Andrea
DOI: 10.1016/j.jbi.2003.11.002
发表时间: 2003-12-01
影响因子: 4.5
作者:
Bodenreider, O;McCray, AT
通讯作者: McCray, AT
DOI: 10.1016/j.jbi.2012.01.010
发表时间: 2012-06-01
影响因子: 4.5
作者:
Chapman, Wendy W.;Savova, Guergana K.;Crowley, Rebecca
通讯作者: Crowley, Rebecca
DOI: 10.1162/0891201053630264
发表时间: 2005-03-01
影响因子: 9.3
作者:
Palmer, M;Kingsbury, P;Gildeafi, D
通讯作者: Gildeafi, D