Classifying protein-protein interaction articles using word and syntactic features.

Classifying protein-protein interaction articles using word and syntactic features.
复制标题

使用单词和句法特征对蛋白质-蛋白质相互作用文章进行分类。

DOI:
10.1186/1471-2105-12-s8-s9
复制
发表时间:
2011-10-03
期刊:
影响因子:
3
通讯作者:
Wilbur WJ
Wilbur WJ
中科院分区:
生物学4区
文献类型:
--
作者:
Kim S;Wilbur WJ

文献摘要

被引文献

相似文献

从文献中识别蛋白质-蛋白质相互作用(PPI)是挖掘单个蛋白质功能及其生物网络的重要步骤。由于已知PPI在文本中具有独特的模式,因此机器学习方法已成功应用于挖掘这些模式。然而,PPI描述的复杂性质使得提取过程困难。我们的方法利用词和句法特征,有效地捕捉PPI模式从生物医学文献。该方法首先通过优先级模型自动识别基因名称,然后使用依赖分析器提取语法关系。一个具有Huber损失函数的大间隔分类器从提取的特征中学习,并使用该数据驱动模型预测未知文章。对于BioCreative III ACT评价,我们的正式运行通过获得最大89.15%准确度、61.42% F1评分、0.55306 MCC评分和67.98% AUC iP/R评分而排名靠前。即使问题仍然存在,利用语法信息进行文章级过滤有助于提高PPI排名性能。拟议的系统是一个修订以前开发的算法在我们的小组的ACT评估。我们的方法是有价值的,在显示如何使用语法关系PPI文章过滤,特别是在有限的训练语料库。虽然目前的性能是远远不能令人满意的注释工具,它已经是有用的PPI文章搜索引擎,因为用户主要集中在高排名的结果。
Identifying protein-protein interactions (PPIs) from literature is an important step in mining the function of individual proteins as well as their biological network. Since it is known that PPIs have distinctive patterns in text, machine learning approaches have been successfully applied to mine these patterns. However, the complex nature of PPI description makes the extraction process difficult. Our approach utilizes both word and syntactic features to effectively capture PPI patterns from biomedical literature. The proposed method automatically identifies gene names by a Priority Model, then extracts grammar relations using a dependency parser. A large margin classifier with Huber loss function learns from the extracted features, and unknown articles are predicted using this data-driven model. For the BioCreative III ACT evaluation, our official runs were ranked in top positions by obtaining maximum 89.15% accuracy, 61.42% F1 score, 0.55306 MCC score, and 67.98% AUC iP/R score. Even though problems still remain, utilizing syntactic information for article-level filtering helps improve PPI ranking performance. The proposed system is a revision of previously developed algorithms in our group for the ACT evaluation. Our approach is valuable in showing how to use grammatical relations for PPI article filtering, in particular, with a limited training corpus. While current performance is far from satisfactory as an annotation tool, it is already useful for a PPI article search engine since users are mainly focused on highly-ranked results.