A Method of Extracting Sentences Containing Protein Function Information from Articles by Iterative Learning with Feature Update

A Method of Extracting Sentences Containing Protein Function Information from Articles by Iterative Learning with Feature Update
复制标题

DOI:
10.1007/978-3-642-38342-7_8
复制
发表时间:
2012-07
期刊:
--
影响因子:
--
通讯作者:
Kazunori Miyanishi;T. Ohkawa
Kazunori Miyanishi;T. Ohkawa
中科院分区:
其他
文献类型:
--
作者:
Kazunori Miyanishi;T. Ohkawa

文献摘要

相似文献

蛋白质是生命系统中重要的大分子,在几乎所有生物过程中发挥各种功能。许多科学文章都报道了蛋白质功能信息。从文章中提取功能信息对于药物发现、生命现象的理解等很有用。然而,从大量文章中手动提取功能信息是不可行的。在本文中,我们提出了一种通过特征更新迭代学习来提取包含蛋白质功能信息的句子的方法。在该方法中,我们使用分类器来区分包含功能信息的句子和其他句子,并引入半自动过程,其中根据用户对先前分类结果的反馈重建新的分类器。在以12篇文章作为反馈数据的实验中,证实了通过迭代学习改进了F-measure,而没有得到反馈的负面影响。
Proteins are important macromolecules in living systems and serve various functions in almost all biological processes. Protein function information is reported in many scientific articles. Extraction of the function information from the articles is useful for drug discovery, understanding of life phenomenon, and so on. However, it is infeasible to extract the function information manually from a number of articles. In this paper, we propose a method of extracting sentences containing protein function information by iterative learning with feature update. In this method, we use a classifier in order to distinguish the sentences containing the function information from the other sentences, and introduce a semi-automatic procedure, in which a new classifier is reconstructed based on the user’s feedback for the previous classified results. In the experiment with twelve articles as feedback data, it was confirmed that F-measure was improved by iterating learning without getting the negative effect of the feedback.