The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full text.

The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full text.
复制标题

DOI:
10.1186/1471-2105-12-s8-s3
复制
发表时间:
2011-10-03
期刊:
影响因子:
3
通讯作者:
Valencia A
Valencia A
中科院分区:
生物学4区
文献类型:
--
作者:
Krallinger M;Vazquez M;Leitner F;Salgado D;Chatr-Aryamontri A;Winter A;Perfetto L;Briganti L;Licata L;Iannuccelli M;Castagnoli L;Cesareni G;Tyers M;Schneider G;Rinaldi F;Leaman R;Gonzalez G;Matos S;Kim S;Wilbur WJ;Rocha L;Shatkay H;Tendulkar AV;Agarwal S;Liu F;Wang X;Rak R;Noto K;Elkan C;Lu Z;Dogan RI;Fontaine JF;Andrade-Navarro MA;Valencia A

文献摘要

被引文献

相似文献

确定生物医学文本挖掘系统的有用性需要现实的任务定义和没有人为约束的数据选择标准,测量超出传统指标的性能方面。BioCreative III蛋白质-蛋白质相互作用(PPI)任务就是出于这样的考虑,试图解决包括最终用户如何监督生成的输出等方面的问题,例如,通过提供排名结果、人为解释的文本证据或通过使用自动化系统测量节省的时间。文章分类任务(ACT)解决了描述ppi等复杂生物事件的文章的检测问题,参与者被要求使用工具来检测描述ppi的摘要。因此,提供了BCIII-ACT语料库,其中包括一个训练、开发和测试集,由领域专家手动标记并记录人类分类时间的12,000多篇PPI相关和不相关的PubMed摘要。交互方法任务(IMT)超越了摘要,需要挖掘3500多篇全文文章和用于检测其中报告的ppi的交互检测方法本体概念之间的关联。共有11个小组至少参与了两个PPI任务中的一个(10个在ACT中,8个在IMT中),共有62人作为参与者或准备数据集/评估这些任务。对于每个任务,每个团队可以通过BioCreative Meta-Server提交5次离线运行和另外5次在线运行。在提交ACT的52次测试中,马修相关系数(MCC)得分最高为0.55,准确率为89%,最佳AUC iP/R为68%。大多数ACT团队探索机器学习方法,其中一些还使用词汇资源,如MeSH术语、PSI-MI概念或特定的动词和名词列表,一些集成了NER方法。对于IMT,通过将系统与管理员从BioGRID和MINT数据库中手动生成的注释进行比较,总共评估了42次运行。各批次的最高AUC iP/R为53%,最佳MCC得分为0.55。在具有可接受召回(高于35%)的竞争系统的情况下,宏观平均精度范围在50%到80%之间,最大F-Score为55%。BioCreative III的ACT任务结果表明,对反映真实类别不平衡的大型不平衡文章集进行分类仍然具有挑战性。然而,报告相关文章排名列表以供人工选择的文本挖掘工具,与未排名的结果相比,可能会将识别一半相关文章所需的时间减少到不到1/4。检测全文文章与交互检测方法PSI-MI术语(IMT)之间的关联比预期的要困难。这是由于方法术语提及的可变性,以PDF文件形式提供的文章的预处理导致的错误,以及本体论中遇到的方法术语概念的异质性和不同粒度。然而,将参与者开发的复杂技术与来自人类解释的文章的支持证据字符串相结合,可能会产生生物注释工作流程的实用模块。
Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows.