Using LMF to Shape a Lexicon for the Biomedical Domain

Using LMF to Shape a Lexicon for the Biomedical Domain
复制标题

使用 LMF 构建生物医学领域词典

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
N. Calzolari
N. Calzolari
中科院分区:
--
文献类型:
--
作者:
M. Monachini;Valeria Quochi;R. Gratta;N. Calzolari

文献摘要

被引文献

相似文献

本文描述了在FP 6项目BootStrep框架下的BioLexicon的设计、实现和人口。BioLexicon(BL)是一个专为生物领域文本挖掘而设计的词汇资源。它被认为是为了满足领域需求和即将到来的词汇表示ISO标准。数据模型和数据类别符合ISO词汇标记框架和数据类别注册表。BioLexicon整合了词汇和术语的功能:从现有资源中提取的术语条目(和变体)富含语言特征,包括从文本中提取的次范畴化和谓词-论元信息。这是一种可扩展的资源。此外,词汇条目将与项目的本体资源BioOntology中的概念保持一致。BL实现是一个可扩展的关系数据库,具有自动填充过程。Population依赖于一个专用的输入数据结构,允许上传术语及其语言属性,并在数据库中“推拉”它们。《生物词典》指出,最先进的技术已经足够成熟,可以在这一领域建立一个标准。由于符合词汇标准,BioLexicon具有互操作性和可移植性。
This paper describes the design, implementation and population of the BioLexicon in the framework of BootStrep, an FP6 project. The BioLexicon (BL) is a lexical resource designed for text mining in the bio-domain. It has been conceived to meet both domain requirements and upcoming ISO standards for lexical representation. The data model and data categories are compliant to the ISO Lexical Markup Framework and the Data Category Registry. The BioLexicon integrates features of lexicons and terminologies: term entries (and variants) derived from existing resources are enriched with linguistic features, including subcategorization and predicate-argument information, extracted from texts. Thus, it is an extendable resource. Furthermore, the lexical entries will be aligned to concepts in the BioOntology, the ontological resource of the project. The BL implementation is an extensible relational database with automatic population procedures. Population relies on a dedicated input data structure allowing to upload terms and their linguistic properties and “pull-and-push” them in the database. The BioLexicon teaches that the state-of-the-art is mature enough to aim at setting up a standard in this domain. Being conformant to lexical standards, the BioLexicon is interoperable and portable to other areas.