Benchmarking of the 2010 BioCreative Challenge III text-mining competition by the BioGRID and MINT interaction databases.

Benchmarking of the 2010 BioCreative Challenge III text-mining competition by the BioGRID and MINT interaction databases.
复制标题

DOI:
10.1186/1471-2105-12-s8-s8
复制
发表时间:
2011-10-03
期刊:
影响因子:
3
通讯作者:
Tyers M
Tyers M
中科院分区:
生物学4区
文献类型:
--
作者:
Chatr-Aryamontri A;Winter A;Perfetto L;Briganti L;Licata L;Iannuccelli M;Castagnoli L;Cesareni G;Tyers M

文献摘要

被引文献

相似文献

主要生物医学文献中发表的大量数据对单个数据元素的自动提取和编纂提出了挑战。仅依靠专家管理员手动提取的生物数据库无法全面注释分散在整个生物医学文献中的信息。基于自然语言处理(NLP)系统的高效工具的开发对于相关出版物的选择、数据属性的识别和部分自动化注释是必不可少的。生物创意2010挑战III的任务之一是致力于评估用于识别蛋白质-蛋白质相互作用(PPI)数据的整理和提取的文章的NLP系统。2010年生物创意大赛有三个任务:基因规范化、文章分类和相互作用方法识别。BioGRID和MINT蛋白互作数据库都参与了基因归一化测试出版集的生成,注释了文章分类的开发和测试集,策划了互作方法分类的测试集。这些测试数据集作为评估数据提取算法的金标准。开发高效的PPI数据提取工具是实现生物医学文献全面管理的必要步骤。NLP系统首先可以通过精炼包含PPI数据的候选出版物列表来促进专家管理;更有野心的是,NLP方法可能能够直接从全文文章中提取相关信息,供专家策展人快速检查。生物数据库和自然语言处理系统开发人员之间的密切合作将继续促进这两个学科的长期目标。
The vast amount of data published in the primary biomedical literature represents a challenge for the automated extraction and codification of individual data elements. Biological databases that rely solely on manual extraction by expert curators are unable to comprehensively annotate the information dispersed across the entire biomedical literature. The development of efficient tools based on natural language processing (NLP) systems is essential for the selection of relevant publications, identification of data attributes and partially automated annotation. One of the tasks of the Biocreative 2010 Challenge III was devoted to the evaluation of NLP systems developed to identify articles for curation and extraction of protein-protein interaction (PPI) data. The Biocreative 2010 competition addressed three tasks: gene normalization, article classification and interaction method identification. The BioGRID and MINT protein interaction databases both participated in the generation of the test publication set for gene normalization, annotated the development and test sets for article classification, and curated the test set for interaction method classification. These test datasets served as a gold standard for the evaluation of data extraction algorithms. The development of efficient tools for extraction of PPI data is a necessary step to achieve full curation of the biomedical literature. NLP systems can in the first instance facilitate expert curation by refining the list of candidate publications that contain PPI data; more ambitiously, NLP approaches may be able to directly extract relevant information from full-text articles for rapid inspection by expert curators. Close collaboration between biological databases and NLP systems developers will continue to facilitate the long-term objectives of both disciplines.