All-paths graph kernel for protein-protein interaction extraction with evaluation of cross-corpus learning.

All-paths graph kernel for protein-protein interaction extraction with evaluation of cross-corpus learning.
复制标题

通过评估跨核心学习,全路径图核用于蛋白质 - 蛋白质相互作用提取。

DOI:
10.1186/1471-2105-9-s11-s2
复制
发表时间:
2008-11-19
期刊:
影响因子:
3
通讯作者:
Salakoski, Tapio
Salakoski, Tapio
中科院分区:
生物学4区
文献类型:
--
作者:
Airola, Antti;Pyysalo, Sampo;Bjoerne, Jari;Pahikkala, Tapio;Ginter, Filip;Salakoski, Tapio

文献摘要

被引文献

相似文献

蛋白质相互作用信息的自动提取是生物医学文本挖掘中一个重要的研究课题。我们提出了一个基于图形内核的方法来完成这项任务。与早期的PPI提取方法相比,引入的所有路径图内核能够利用表示句子结构的完整、通用依赖图。我们评估了所提出的方法在五个公开的PPI语料库,提供了最全面的评估基于机器学习的PPI提取系统。我们还对不同资源上的训练和测试效果进行了详细评估,深入了解了应用系统时所面临的挑战,这些挑战超出了它所训练的数据。我们的方法显示出在可比评估方面达到了最先进的性能,在AImed语料库上的F分数为56.4,AUC为84.8。我们表明,图核方法在PPI提取中的性能达到了最先进的水平,并注意到提取复杂交互作用任务的可能扩展。跨语料库的结果提供了进一步的了解如何学习推广超越个人语料库。此外,我们确定了几个陷阱,可以使PPI提取系统的评价无与伦比,甚至无效。这些问题包括不正确的交叉验证策略以及与比较不同评估资源上的F分数结果相关的问题。提供了避免这些陷阱的建议。
Automated extraction of protein-protein interactions (PPI) is an important and widely studied task in biomedical text mining. We propose a graph kernel based approach for this task. In contrast to earlier approaches to PPI extraction, the introduced all-paths graph kernel has the capability to make use of full, general dependency graphs representing the sentence structure. We evaluate the proposed method on five publicly available PPI corpora, providing the most comprehensive evaluation done for a machine learning based PPI-extraction system. We additionally perform a detailed evaluation of the effects of training and testing on different resources, providing insight into the challenges involved in applying a system beyond the data it was trained on. Our method is shown to achieve state-of-the-art performance with respect to comparable evaluations, with 56.4 F-score and 84.8 AUC on the AImed corpus. We show that the graph kernel approach performs on state-of-the-art level in PPI extraction, and note the possible extension to the task of extracting complex interactions. Cross-corpus results provide further insight into how the learning generalizes beyond individual corpora. Further, we identify several pitfalls that can make evaluations of PPI-extraction systems incomparable, or even invalid. These include incorrect cross-validation strategies and problems related to comparing F-score results achieved on different evaluation resources. Recommendations for avoiding these pitfalls are provided.