Building a semantically annotated corpus of clinical texts

Building a semantically annotated corpus of clinical texts
复制标题

DOI:
10.1016/j.jbi.2008.12.013
复制
发表时间:
2009-10-01
影响因子:
4.5
通讯作者:
Setzer, Andrea
Setzer, Andrea
中科院分区:
医学3区
文献类型:
--
作者:
Roberts, Angus;Gaizauskas, Robert;Setzer, Andrea

文献摘要

被引文献

相似文献

在本文中,我们描述了一个语义注释的临床文本语料库的构建,用于开发和评估系统,用于自动从患者记录的文本成分中提取临床重要信息。本文详细介绍了从20000份癌症患者病历中抽取文本材料,开发语义标注方案,标注方法,在最终语料库中标注的分布,以及使用语料库开发自适应信息提取系统。由此产生的语料库是迄今为止为临床文本处理构建的最丰富的语义注释资源,其价值已通过其在开发有效信息提取系统中的使用得到证明。我们的语料库构建和注释方法的详细介绍将对其他寻求在生物医学领域建立高质量语义注释语料库的人有价值。(C) 2009爱思唯尔公司版权所有。
In this paper, we describe the construction of a semantically annotated corpus of clinical texts for use in the development and evaluation of systems for automatically extracting clinically significant information from the textual component of patient records. The paper details the sampling of textual material from a collection of 20,000 cancer patient records, the development of a semantic annotation scheme, the annotation methodology, the distribution of annotations in the final corpus, and the use of the corpus for development of an adaptive information extraction system. The resulting corpus is the most richly semantically annotated resource for clinical text processing built to date, whose value has been demonstrated through its use in developing an effective information extraction system. The detailed presentation of our corpus construction and annotation methodology will be of value to others seeking to build high-quality semantically annotated corpora in biomedical domains. (C) 2009 Elsevier Inc. All rights reserved.