Extraction of Protein-Protein Interaction from Scientific Articles by Predicting Dominant Keywords.

Extraction of Protein-Protein Interaction from Scientific Articles by Predicting Dominant Keywords.
复制标题

DOI:
10.1155/2015/928531
复制
发表时间:
2015
影响因子:
--
通讯作者:
Ohkawa T
Ohkawa T
中科院分区:
生物学3区
文献类型:
--
作者:
Koyabu S;Phan TT;Ohkawa T

文献摘要

被引文献

相似文献

对于从科学文章中自动提取蛋白质-蛋白质相互作用信息,机器学习方法是有用的。分类器从使用几个特征表示的训练数据中生成,以确定每个句子中的蛋白质对是否具有相互作用。“bind”或“interaction”等与交互直接相关的特定关键字对于训练分类器起着重要的作用。我们称其为影响分类器能力的主导关键字。虽然识别占主导地位的关键字很重要,但关键字是否占主导地位取决于它出现的上下文。因此,我们提出了一种预测关键字在每个实例中是否占主导地位的方法。在该方法中,初步假定产生不平衡分类结果的关键字为主导关键字。然后从具有和不具有假设的主导关键字的实例中分别训练分类器。根据生成的分类器的分类结果评估假设的主导关键字的有效性。根据评估结果对假设进行更新。重复这个过程可以提高主要关键字的预测精度。使用5个语料库的实验结果表明,该方法具有优势关键词预测的有效性。
For the automatic extraction of protein-protein interaction information from scientific articles, a machine learning approach is useful. The classifier is generated from training data represented using several features to decide whether a protein pair in each sentence has an interaction. Such a specific keyword that is directly related to interaction as “bind” or “interact” plays an important role for training classifiers. We call it a dominant keyword that affects the capability of the classifier. Although it is important to identify the dominant keywords, whether a keyword is dominant depends on the context in which it occurs. Therefore, we propose a method for predicting whether a keyword is dominant for each instance. In this method, a keyword that derives imbalanced classification results is tentatively assumed to be a dominant keyword initially. Then the classifiers are separately trained from the instance with and without the assumed dominant keywords. The validity of the assumed dominant keyword is evaluated based on the classification results of the generated classifiers. The assumption is updated by the evaluation result. Repeating this process increases the prediction accuracy of the dominant keyword. Our experimental results using five corpora show the effectiveness of our proposed method with dominant keyword prediction.