A Large Scale Terminology Resource for Biomedical Text Processing

A Large Scale Terminology Resource for Biomedical Text Processing
复制标题

用于生物医学文本处理的大规模术语资源

DOI:
--
复制
发表时间:
2004
期刊:
HLT-NAACL 2004
影响因子:
--
通讯作者:
Yikun Guo
Yikun Guo
中科院分区:
--
文献类型:
--
作者:
H. Harkema;R. Gaizauskas;Mark Hepple;A. Roberts;Ian Roberts;Neil Davis;Yikun Guo

文献摘要

被引文献

相似文献

在本文中,我们讨论的设计,实现和使用的Termino,一个大规模的术语资源的文本处理。术语处理是语言处理应用中一项困难但不可避免的任务,例如技术领域的信息提取。必须存储关于大量术语的复杂、异构的信息。同时,术语识别必须在现实时代进行。Termino试图通过维护一个灵活的、可扩展的关系数据库来存储术语信息,并从该数据库编译有限状态机来进行术语查找,从而调和这种紧张关系。虽然Termino是为生物医学应用开发的,但其通用设计允许其用于任何领域的术语处理。
In this paper we discuss the design, implementation, and use of Termino, a large scale terminological resource for text processing. Dealing with terminology is a difficult but unavoidable task for language processing applications, such as Information Extraction in technical domains. Complex, heterogeneous information must be stored about large numbers of terms. At the same time term recognition must be performed in realistic times. Termino attempts to reconcile this tension by maintaining a flexible, extensible relational database for storing terminological information and compiling finite state machines from this database to do term lookup. While Termino has been developed for biomedical applications, its general design allows it to be used for term processing in any domain.