GENIA corpus-a semantically annotated corpus for bio-textmining

GENIA corpus-a semantically annotated corpus for bio-textmining
复制标题

DOI:
10.1093/bioinformatics/btg1023
复制
发表时间:
2003-07-01
期刊:
影响因子:
5.8
通讯作者:
Tsujii, J.
Tsujii, J.
中科院分区:
生物学3区
文献类型:
--
作者:
Kim, J-D;Ohta, T.;Tsujii, J.

文献摘要

被引文献

相似文献

动机:自然语言处理(NLP)方法被认为是有用的,以提高从生物文献中挖掘文本的潜力。然而,缺乏一个广泛注释的语料库,这方面的文献,导致应用NLP技术的主要瓶颈。GENIA语料库正在开发中,为自然语言处理技术在生物文本挖掘中的应用提供参考资料。结果:GENIA语料库3.0版已发布,包含2000篇MEDLINE摘要,40多万字,近10万条生物学术语注释。
Motivation: Natural language processing (NLP) methods are regarded as being useful to raise the potential of text mining from biological literature. The lack of an extensively annotated corpus of this literature, however, causes a major bottleneck for applying NLP techniques. GENIA corpus is being developed to provide reference materials to let NLP techniques work for bio-textmining.Results: GENIA corpus version 3.0 consisting of 2000 MEDLINE abstracts has been released with more than 400 000 words and almost 100 000 annotations for biological terms.