A lexicon for biology and bioinformatics: the BOOTStrep experience.

A lexicon for biology and bioinformatics: the BOOTStrep experience.
复制标题

生物学和生物信息学词典:BOOTStrep 体验。

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
N. Calzolari
N. Calzolari
中科院分区:
--
文献类型:
--
作者:
Valeria Quochi;M. Monachini;R. Gratta;N. Calzolari

文献摘要

被引文献

相似文献

本文描述了在一个正在进行的欧洲项目中开发的生物学和生物信息学词汇资源(BioLexicon)的设计、实施和填充。该项目的目标是基于文本的知识收集,以支持生物医学领域的信息提取和文本挖掘。 BioLexicon 是一种大型词汇术语资源,将不同信息类型编码在一个集成资源中。在资源的设计中,我们遵循 ISO/DIS 24613 词汇标记框架标准,确保编码信息的可重用性以及数据和架构的轻松交换。该资源的设计还考虑了我们的文本挖掘合作伙伴的需求,他们自动从文本中提取句法和语义信息并将其提供给词典。本贡献首先详细描述了 BioLexicon 的模型的三个主要层次:词法、句法和语义;然后,简要描述了模型的数据库实现以及项目中遵循的人口策略,并举例说明。事实上,BioLexicon 数据库配备了基于通用交换 XML 格式的自动上传程序,这保证了词典可以正确填充来自不同来源的数据。
This paper describes the design, implementation and population of a lexical resource for biology and bioinformatics (the BioLexicon) developed within an ongoing European project. The aim of this project is text-based knowledge harvesting for support to information extraction and text mining in the biomedical domain. The BioLexicon is a large-scale lexical-terminological resource encoding different information types in one single integrated resource. In the design of the resource we follow the ISO/DIS 24613 Lexical Mark-up Framework standard, which ensures reusability of the information encoded and easy exchange of both data and architecture. The design of the resource also takes into account the needs of our text mining partners who automatically extract syntactic and semantic information from texts and feed it to the lexicon. The present contribution first describes in detail the model of the BioLexicon along its three main layers: morphology, syntax and semantics; then, it briefly describes the database implementation of the model and the population strategy followed within the project, together with an example. The BioLexicon database in fact comes equipped with automatic uploading procedures based on a common exchange XML format, which guarantees that the lexicon can be properly populated with data coming from different sources.