GENETAG: a tagged corpus for gene/protein named entity recognition.

GENETAG: a tagged corpus for gene/protein named entity recognition.
复制标题

Genetag:一种名为实体识别的基因/蛋白质的标记语料库。

DOI:
10.1186/1471-2105-6-s1-s3
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Wilbur, WJ
Wilbur, WJ
中科院分区:
生物学4区
文献类型:
--
作者:
Tanabe, L;Xie, N;Thom, LH;Matten, W;Wilbur, WJ

文献摘要

被引文献

相似文献

命名实体识别(NER)是生物医学文献文本挖掘的重要第一步。如果没有标准化的测试语料库,就不可能评估生物医学 NER 系统的性能。由于基因/蛋白质名称的复杂性,对这样的基因/蛋白质名称NER的语料库进行注释是一个困难的过程。我们描述了 GENETAG 的构建和注释,这是一个包含 20K MEDLINE® 基因/蛋白质 NER 句子的语料库。 BioCreAtIvE 任务 1A 竞赛使用了 15K GENETAG 句子。为了确保语料库的异质性,首先对 MEDLINE 句子与已知基因名称的文档的术语相似度进行评分,并随机选择 10K 高分和 10K 低分句子。原始的 20K 句子通过基因/蛋白质名称标记器运行,并手动修改结果以反映受特异性约束的基因/蛋白质名称的广泛定义,该规则要求标记的实体引用特定实体。 GENETAG 中的每个句子都用其包含的基因/蛋白质名称的可接受替代品进行注释,从而允许与语义约束进行部分匹配。语义约束是要求标记实体在句子上下文中包含其真实含义的规则。与无限制的部分匹配相比,应用这些约束可以更有意义地衡量 NER 系统的性能。 GENETAG的注释需要注释者进行复杂的手动判断,这阻碍了标记的一致性。数据被预先分割成单词,以提供支持系统响应与“黄金标准”比较的索引。然而,基于字符的索引比基于单词的索引更稳健。 GENETAG 训练、测试和 Round1 数据及辅助程序可在 上免费获取。 GENETAG-05 的新版本将于今年晚些时候发布。
Named entity recognition (NER) is an important first step for text mining the biomedical literature. Evaluating the performance of biomedical NER systems is impossible without a standardized test corpus. The annotation of such a corpus for gene/protein name NER is a difficult process due to the complexity of gene/protein names. We describe the construction and annotation of GENETAG, a corpus of 20K MEDLINE® sentences for gene/protein NER. 15K GENETAG sentences were used for the BioCreAtIvE Task 1A Competition. To ensure heterogeneity of the corpus, MEDLINE sentences were first scored for term similarity to documents with known gene names, and 10K high- and 10K low-scoring sentences were chosen at random. The original 20K sentences were run through a gene/protein name tagger, and the results were modified manually to reflect a wide definition of gene/protein names subject to a specificity constraint, a rule that required the tagged entities to refer to specific entities. Each sentence in GENETAG was annotated with acceptable alternatives to the gene/protein names it contained, allowing for partial matching with semantic constraints. Semantic constraints are rules requiring the tagged entity to contain its true meaning in the sentence context. Application of these constraints results in a more meaningful measure of the performance of an NER system than unrestricted partial matching. The annotation of GENETAG required intricate manual judgments by annotators which hindered tagging consistency. The data were pre-segmented into words, to provide indices supporting comparison of system responses to the "gold standard". However, character-based indices would have been more robust than word-based indices. GENETAG Train, Test and Round1 data and ancillary programs are freely available at . A newer version of GENETAG-05, will be released later this year.