Overview of BioCreAtIvE: critical assessment of information extraction for biology.

Overview of BioCreAtIvE: critical assessment of information extraction for biology.
复制标题

生物公约概述:生物学信息提取的批判性评估。

DOI:
10.1186/1471-2105-6-s1-s1
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Valencia A
Valencia A
中科院分区:
生物学4区
文献类型:
--
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A

文献摘要

被引文献

相似文献

第一个 BioCreAtIvE 挑战(生物学信息提取的批判性评估)的目标是提供一组常见的评估任务,以评估应用于生物学问题的文本挖掘的最新技术水平。结果于 2004 年 3 月 28 日至 31 日在西班牙格拉纳达举行的研讨会上公布。BMC 生物信息学增刊中收集的题为“分子生物学中文本挖掘方法的批判性评估”的文章描述了 BioCreAtIvE 任务、系统、结果及其独立评估。 BioCreAtIvE 专注于两项任务。第一个涉及从文本中提取基因或蛋白质名称,并将其映射到三个模式生物数据库(果蝇、小鼠、酵母)的标准化基因标识符。第二项任务解决了功能注释的问题,要求系统在给定全文文章的情况下识别支持特定蛋白质的基因本体注释的特定文本段落。首届 BioCreAtIvE 评估获得了高水平的国际参与(来自 10 个国家的 27 个小组)。该评估为基本任务(基因名称查找和标准化)提供了最先进的性能结果,其中最好的系统达到了 80% 的平衡精确度/召回率或更好,这可能使它们适合生物学中的实际应用。高级任务(自由文本的功能注释)的结果明显较低,这表明当前需要知识外推和解释的文本挖掘方法的局限性。此外,BioCreAtIvE 的一个重要贡献是为这两项任务创建和发布了训练和测试数据集。本期特刊共有 22 篇文章,其中 6 篇文章提供了数据集结果或数据质量的分析,其中包括针对任务 2 中使用的测试集的新颖的注释者间一致性评估。
The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28–31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.