Overview of the BioCreative III Workshop.

Overview of the BioCreative III Workshop.
复制标题

DOI:
10.1186/1471-2105-12-s8-s1
复制
发表时间:
2011-10-03
期刊:
影响因子:
3
通讯作者:
Wu CH
Wu CH
中科院分区:
生物学4区
文献类型:
--
作者:
Arighi CN;Lu Z;Krallinger M;Cohen KB;Wilbur WJ;Valencia A;Hirschman L;Wu CH

文献摘要

被引文献

相似文献

BioCreative研讨会的总体目标是促进文本挖掘和文本处理工具的开发,这些工具对生物科学的研究人员和数据库管理员社区有用。为此,2004年举办了BioCreative I,2007年举办了BioCreative II,2009年举办了BioCreative II.5。每个研讨会都涉及人工注释的测试数据,用于应用于生物医学文献的文本挖掘的几个基本任务。参加工作坊的人士获邀参与设计软件系统,以自动完成工作,并根据他们的表现给予分数。这些讲习班的成果在几个方面使社区受益。它们1)为目前解决特定问题的最有效方法提供了证据; 2)揭示了解决这些问题的最新技术水平; 3)并提供了黄金标准数据和数据结果,通过这些数据可以衡量未来的进步。本特刊包含BioCreative III三项任务的概述论文。BioCreative III研讨会于2010年9月举行,延续了对生物学中有效文本挖掘的几个基本任务进行挑战性评估的传统,包括基因标准化(GN)任务和两个蛋白质-蛋白质相互作用(PPI)任务。讲习班共涉及23个小组的工作。13个团队参与了GN任务,该任务要求在全文论文中为所有命名的基因分配RzGene ID,而不向系统提供任何物种信息。10个团队参与了PPI文章分类任务(ACT),要求系统将PubMed®记录分类和排序为属于具有或不具有“PPI相关”信息的文章。八个团队参加了PPI交互方法任务(IMT),其中系统被给予全文文档,并被要求提取用于建立PPI的实验方法和支持每个方法的文本片段。为每项任务编制了黄金标准数据,参与者竞相开发自动执行任务的系统。BioCreative III还引入了一个新的交互式任务(IAT),作为演示任务运行。目标是开发一个交互式系统,以方便用户注释文章中出现的所有基因的唯一数据库标识符。该任务包括按重要性对基因进行排序(优选地基于所描述的关于基因的实验信息的量)。还有一个可选任务,帮助用户找到与给定基因最相关的文章。对于BioCreative III,组建了一个用户咨询小组(UAG),并在以下方面发挥了重要作用:1)为GN任务制作一些黄金标准注释,2)批评IAT系统,3)为未来更严格的IAT系统评价提供指导。六个小组参加了IAT演示任务,并收到了UAG小组对其系统的反馈。除了在一般网络和生产者价格指数任务中进行创新,使其更加现实和实用,并引入了机构间评估测试任务之外,还开始讨论共同体数据标准,以促进互操作性,并讨论用户要求和评价标准,以解决系统的效用和可用性问题。在本文中,我们简要介绍了生物创意工作坊的历史,以及它们与生物学中其他文本挖掘比赛的关系。接下来是BioCreative III中GN、PPI和IAT三个任务的概要,以及GN和PPI任务的最佳参与者表现的数字。这些结果进行了讨论,并与以前的BioCreative研讨会的结果进行了比较,我们得出结论,在现实环境中,GN,PPI-ACT和PPI-IMT的最佳性能系统不足以实现全自动使用。这为交互式系统的重要性提供了证据,我们提出了我们的愿景,如何最好地构建一个交互式系统的GN或PPI一样的任务,在剩余的文件。
The overall goal of the BioCreative Workshops is to promote the development of text mining and text processing tools which are useful to the communities of researchers and database curators in the biological sciences. To this end BioCreative I was held in 2004, BioCreative II in 2007, and BioCreative II.5 in 2009. Each of these workshops involved humanly annotated test data for several basic tasks in text mining applied to the biomedical literature. Participants in the workshops were invited to compete in the tasks by constructing software systems to perform the tasks automatically and were given scores based on their performance. The results of these workshops have benefited the community in several ways. They have 1) provided evidence for the most effective methods currently available to solve specific problems; 2) revealed the current state of the art for performance on those problems; 3) and provided gold standard data and results on that data by which future advances can be gauged. This special issue contains overview papers for the three tasks of BioCreative III. The BioCreative III Workshop was held in September of 2010 and continued the tradition of a challenge evaluation on several tasks judged basic to effective text mining in biology, including a gene normalization (GN) task and two protein-protein interaction (PPI) tasks. In total the Workshop involved the work of twenty-three teams. Thirteen teams participated in the GN task which required the assignment of EntrezGene IDs to all named genes in full text papers without any species information being provided to a system. Ten teams participated in the PPI article classification task (ACT) requiring a system to classify and rank a PubMed® record as belonging to an article either having or not having “PPI relevant” information. Eight teams participated in the PPI interaction method task (IMT) where systems were given full text documents and were required to extract the experimental methods used to establish PPIs and a text segment supporting each such method. Gold standard data was compiled for each of these tasks and participants competed in developing systems to perform the tasks automatically. BioCreative III also introduced a new interactive task (IAT), run as a demonstration task. The goal was to develop an interactive system to facilitate a user’s annotation of the unique database identifiers for all the genes appearing in an article. This task included ranking genes by importance (based preferably on the amount of described experimental information regarding genes). There was also an optional task to assist the user in finding the most relevant articles about a given gene. For BioCreative III, a user advisory group (UAG) was assembled and played an important role 1) in producing some of the gold standard annotations for the GN task, 2) in critiquing IAT systems, and 3) in providing guidance for a future more rigorous evaluation of IAT systems. Six teams participated in the IAT demonstration task and received feedback on their systems from the UAG group. Besides innovations in the GN and PPI tasks making them more realistic and practical and the introduction of the IAT task, discussions were begun on community data standards to promote interoperability and on user requirements and evaluation metrics to address utility and usability of systems. In this paper we give a brief history of the BioCreative Workshops and how they relate to other text mining competitions in biology. This is followed by a synopsis of the three tasks GN, PPI, and IAT in BioCreative III with figures for best participant performance on the GN and PPI tasks. These results are discussed and compared with results from previous BioCreative Workshops and we conclude that the best performing systems for GN, PPI-ACT and PPI-IMT in realistic settings are not sufficient for fully automatic use. This provides evidence for the importance of interactive systems and we present our vision of how best to construct an interactive system for a GN or PPI like task in the remainder of the paper.