Extraction of regulatory gene/protein networks from Medline

Extraction of regulatory gene/protein networks from Medline
复制标题

DOI:
10.1093/bioinformatics/bti597
复制
发表时间:
2006-03-15
期刊:
影响因子:
5.8
通讯作者:
Bork, P
Bork, P
中科院分区:
生物学3区
文献类型:
--
作者:
Saric, J;Jensen, LJ;Bork, P

文献摘要

被引文献

相似文献

动机:我们以前开发了一种基于规则的方法,用于提取酵母中基因表达调控的信息。生物医学文献,但是,包含其他几个同样重要的监管机制,特别是磷酸化,我们现在扩大了我们的规则为基础的系统也extraction.Results:本文提出了新的结果,从生物医学文本中提取的关系信息。我们已经改进了我们的系统STRING-IE,以捕获新类型的语言结构以及新类型的生物信息[即(去)磷酸化]。精确度保持稳定,召回率略有增加。从近100万PubMed摘要相关的四个模式生物,我们设法提取监管网络和二元磷酸化,包括3319关系块。基因表达和(去)磷酸化关系的准确度分别为83-90%和86-95%。为了实现这一目标,我们利用了一个生物体特定的基因/蛋白质名称的资源大大大于大多数其他生物学相关的信息提取方法中使用的。在GENIA语料库上重新训练词性(POS)标记器时,这些名称被包含在词典中。对于所讨论的域,POS标签的准确率达到96.4%。应该注意的是,这些规则是为酵母开发的,并成功地应用于与其他生物相关的摘要和全文文章,具有相当的准确性。
Motivation: We have previously developed a rule-based approach for extracting information on the regulation of gene expression in yeast. The biomedical literature, however, contains information on several other equally important regulatory mechanisms, in particular phosphorylation, which we now expanded for our rule-based system also to extract.Results: This paper presents new results for extraction of relational information from biomedical text. We have improved our system, STRING-IE, to capture both new types of linguistic constructs as well as new types of biological information [i.e. (de-)phosphorylation]. The precision remains stable with a slight increase in recall. From almost one million PubMed abstracts related to four model organisms, we manage to extract regulatory networks and binary phosphorylations comprising 3319 relation chunks. The accuracy is 83-90% and 86-95% for gene expression and (de-)phosphorylation relations, respectively. To achieve this, we made use of an organism-specific resource of gene/protein names considerably larger than those used in most other biology related information extraction approaches. These names were included in the lexicon when retraining the part-of-speech (POS) tagger on the GENIA corpus. For the domain in question, an accuracy of 96.4% was attained on POS tags. It should be noted that the rules were developed for yeast and successfully applied to both abstracts and full-text articles related to other organisms with comparable accuracy.