BioNLP Shared Task--The Bacteria Track.

BioNLP Shared Task--The Bacteria Track.
复制标题

DOI:
10.1186/1471-2105-13-s11-s3
复制
发表时间:
2012-06-26
期刊:
影响因子:
3
通讯作者:
Nédellec C
Nédellec C
中科院分区:
生物学4区
文献类型:
--
作者:
Bossy R;Jourde J;Manine AP;Veber P;Alphonse E;van de Guchte M;Bessières P;Nédellec C

文献摘要

被引文献

相似文献

我们展示了BioNLP 2011共享任务细菌轨迹,这是第一个完全致力于细菌的信息提取挑战。它包括三项任务,涵盖不同水平的生物学知识。细菌基因更名支持任务旨在提取PubMed摘要中的基因更名和基因同义词。细菌基因交互作用是从单个句子中提取基因/蛋白质交互作用的任务。这些相互作用被分成十个不同的亚类,从而在分子水平上详细说明了遗传规律。最后,细菌生物群任务的重点是教科书文章中提到的细菌的本地化和环境。我们描述了这三个语料库的创建过程,包括文档获取和手动注释,以及用于评估参与者提交的指标。三个团队提交了细菌基因重命名任务;最好的团队获得了87%的F-分数。在细菌基因相互作用任务中,唯一参与者的分数达到了77%的全局F分数,尽管系统效率因子类型而异。三个团队以非常不同的方法提交了细菌生物群任务;最好的团队获得了45%的F-分数。然而,对参与制度效率的详细研究揭示了每个参与制度的优势和劣势。细菌跟踪的三项任务为参与者提供了一个机会来解决信息提取中的广泛问题,包括实体识别、语义分类和共指解析。我们在最有效的系统中发现了共同的趋势:系统地使用语法依存关系和机器学习。然而,细菌生物群任务的独创性鼓励了有趣的新方法和技术的使用,例如术语合成性,其范围比句子更广。
We present the BioNLP 2011 Shared Task Bacteria Track, the first Information Extraction challenge entirely dedicated to bacteria. It includes three tasks that cover different levels of biological knowledge. The Bacteria Gene Renaming supporting task is aimed at extracting gene renaming and gene name synonymy in PubMed abstracts. The Bacteria Gene Interaction is a gene/protein interaction extraction task from individual sentences. The interactions have been categorized into ten different sub-types, thus giving a detailed account of genetic regulations at the molecular level. Finally, the Bacteria Biotopes task focuses on the localization and environment of bacteria mentioned in textbook articles. We describe the process of creation for the three corpora, including document acquisition and manual annotation, as well as the metrics used to evaluate the participants' submissions. Three teams submitted to the Bacteria Gene Renaming task; the best team achieved an F-score of 87%. For the Bacteria Gene Interaction task, the only participant's score had reached a global F-score of 77%, although the system efficiency varies significantly from one sub-type to another. Three teams submitted to the Bacteria Biotopes task with very different approaches; the best team achieved an F-score of 45%. However, the detailed study of the participating systems efficiency reveals the strengths and weaknesses of each participating system. The three tasks of the Bacteria Track offer participants a chance to address a wide range of issues in Information Extraction, including entity recognition, semantic typing and coreference resolution. We found commond trends in the most efficient systems: the systematic use of syntactic dependencies and machine learning. Nevertheless, the originality of the Bacteria Biotopes task encouraged the use of interesting novel methods and techniques, such as term compositionality, scopes wider than the sentence.