Literature mining and database annotation of protein phosphorylation using a rule-based system

Literature mining and database annotation of protein phosphorylation using a rule-based system
复制标题

DOI:
10.1093/bioinformatics/bti390
复制
发表时间:
2005-06-01
期刊:
影响因子:
5.8
通讯作者:
Wu, CH
Wu, CH
中科院分区:
生物学3区
文献类型:
--
作者:
Hu, ZZ;Narayanaswamy, M;Wu, CH

文献摘要

被引文献

相似文献

动机:关于蛋白质磷酸化的大量实验数据被埋葬在快速增长的PubMed文献中。尽管具有巨大的价值,但由于基于文学的策展的费力,此类信息受到数据库的限制。计算文献挖掘有望促进数据库策划。种族:一种基于规则的系统RLIMS-P(基于规则的蛋白质磷酸化文献挖掘系统),用于从Medline摘要中提取蛋白质磷酸化信息。在PIR上开发的注释标记的文献用于评估从摘要中找到磷酸化论文并从摘要中提取磷酸化对象(激酶,底物和位点)的系统。 RLIMS-P的纸质检索获得了91.4和96.4%的精确度和召回率,用于提取底物和地点的97.9和88.0%。将高度召回纸的回忆耦合,以获取信息提取,RLIMS-P促进了蛋白质磷酸化的文献挖掘和数据库注释。
Motivation: A large volume of experimental data on protein phosphorylation is buried in the fast-growing PubMed literature. While of great value, such information is limited in databases owing to the laborious process of literature-based curation. Computational literature mining holds promise to facilitate database curation.Results: A rule-based system, RLIMS-P (Rule-based LIterature Mining System for Protein Phosphorylation), was used to extract protein phosphorylation information from MEDLINE abstracts. An annotation-tagged literature corpus developed at PIR was used to evaluate the system for finding phosphorylation papers and extracting phosphorylation objects (kinases, substrates and sites) from abstracts. RLIMS-P achieved a precision and recall of 91.4 and 96.4% for paper retrieval, and of 97.9 and 88.0% for extraction of substrates and sites. Coupling the high recall for paper retrieval and high precision for information extraction, RLIMS-P facilitates literature mining and database annotation of protein phosphorylation.