Concept annotation in the CRAFT corpus.

Concept annotation in the CRAFT corpus.
复制标题

DOI:
10.1186/1471-2105-13-161
复制
发表时间:
2012-07-09
期刊:
影响因子:
3
通讯作者:
Hunter LE
Hunter LE
中科院分区:
生物学4区
文献类型:
--
作者:
Bada M;Eckert M;Evans D;Garcia K;Shipley K;Sitnikov D;Baumgartner WA Jr;Cohen KB;Verspoor K;Blake JA;Hunter LE

文献摘要

参考文献

被引文献

相似文献

手动注释的语料库对于识别生物医学文本中的概念的自动化方法的培训和评估至关重要。本文介绍了科罗拉多州丰富注释全文 (CRAFT) 语料库的概念注释,该语料库包含 97 篇完整的、开放获取的生物医学期刊文章,这些文章已在语义和句法上进行注释,可作为生物医学自然语言处理 (NLP) 社区的研究资源。 CRAFT 识别了来自九个著名生物医学本体论和术语的几乎所有概念的所有提及:细胞类型本体论、生物兴趣本体论的化学实体、NCBI 分类学、蛋白质本体论、序列本体论、Entrez 基因数据库的条目以及基因本体论的三个子本体论。首次公开发布包括 97 篇文章中 67 篇的注释,为未来的文本挖掘竞赛保留了两组 15 篇文章(之后这些文章也将发布)。概念注释是根据一组指南创建的,这使我们能够始终保持较高的注释者间一致性。由于最初发布的 67 篇文章包含超过 560,000 个标记(全套超过 790,000 个标记),我们的语料库是最大的黄金标准注释生物医学语料库之一。与大多数其他文章不同的是,构成该语料库的期刊文章来自不同的生物医学学科,并进行了整体标记。此外,67 篇文章子集中的概念注释数量接近 100,000 个(整个集合中的概念注释数量超过 140,000 个),概念标记的规模也是同类语料库中最大的。 CRAFT 语料库的概念注释有潜力为 NLP 系统提供高质量的黄金标准,从而显着推进生物医学文本挖掘。语料库、注释指南和其他相关资源可在 http://bionlp-corpora.sourceforge.net/CRAFT/index.shtml 上免费获取。
Manually annotated corpora are critical for the training and evaluation of automated methods to identify concepts in biomedical text. This paper presents the concept annotations of the Colorado Richly Annotated Full-Text (CRAFT) Corpus, a collection of 97 full-length, open-access biomedical journal articles that have been annotated both semantically and syntactically to serve as a research resource for the biomedical natural-language-processing (NLP) community. CRAFT identifies all mentions of nearly all concepts from nine prominent biomedical ontologies and terminologies: the Cell Type Ontology, the Chemical Entities of Biological Interest ontology, the NCBI Taxonomy, the Protein Ontology, the Sequence Ontology, the entries of the Entrez Gene database, and the three subontologies of the Gene Ontology. The first public release includes the annotations for 67 of the 97 articles, reserving two sets of 15 articles for future text-mining competitions (after which these too will be released). Concept annotations were created based on a single set of guidelines, which has enabled us to achieve consistently high interannotator agreement. As the initial 67-article release contains more than 560,000 tokens (and the full set more than 790,000 tokens), our corpus is among the largest gold-standard annotated biomedical corpora. Unlike most others, the journal articles that comprise the corpus are drawn from diverse biomedical disciplines and are marked up in their entirety. Additionally, with a concept-annotation count of nearly 100,000 in the 67-article subset (and more than 140,000 in the full collection), the scale of conceptual markup is also among the largest of comparable corpora. The concept annotations of the CRAFT Corpus have the potential to significantly advance biomedical text mining by providing a high-quality gold standard for NLP systems. The corpus, annotation guidelines, and other associated resources are freely available at http://bionlp-corpora.sourceforge.net/CRAFT/index.shtml.
DOI: 10.1002/cfg.91
发表时间: 2001
影响因子: --
作者:
Blaschke, C;Valencia, A
通讯作者: Valencia, A
DOI: 10.1093/bioinformatics/bth386
发表时间: 2004-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Corney, DPA;Buxton, BF;Jones, DT
通讯作者: Jones, DT
DOI: 10.1186/1471-2105-11-492
发表时间: 2010-09-29
期刊: BMC bioinformatics
影响因子: 3
作者:
Cohen KB;Johnson HL;Verspoor K;Roeder C;Hunter LE
通讯作者: Hunter LE
DOI: 10.1186/gb-2005-6-2-r21
发表时间: 2005
期刊: Genome biology
影响因子: 12.3
作者:
Bard J;Rhee SY;Ashburner M
通讯作者: Ashburner M
DOI: 10.1016/j.jbi.2005.02.009
发表时间: 2005-12-01
影响因子: 4.5
作者:
Coden, AR;Pakhomov, SV;Chute, CG
通讯作者: Chute, CG