IProLINK: an integrated protein resource for literature mining

IProLINK: an integrated protein resource for literature mining
复制标题

DOI:
10.1016/j.compbiolchem.2004.09.010
复制
发表时间:
2004-12-01
影响因子:
3.1
通讯作者:
Wu, CH
Wu, CH
中科院分区:
生物学3区
文献类型:
--
作者:
Hu, ZZ;Mani, I;Wu, CH

文献摘要

被引文献

相似文献

大规模分子序列数据和PubMed科学文献的指数增长促使生物文献挖掘和信息提取方面的积极研究,以促进基因组/蛋白质组注释并提高生物数据库的质量。由于文本挖掘方法的前景,但同时缺乏足够的训练和基准测试数据,蛋白质信息资源(PIR)开发了一个用于蛋白质文献挖掘的资源iProLINK(集成蛋白质文献信息和知识)。由于PIR的工作重点是管理UniProt蛋白质序列数据库,iProLINK的目标是提供可用于书目映射、注释提取等领域的文本挖掘研究的管理数据源。蛋白质命名实体识别和蛋白质本体开发。用于书目映射和注释提取的数据源包括映射的引文(PubMed ID到蛋白质条目和特征线映射)和注释标记的文献语料库。后者包括数百篇摘要和全文文章,其中标注了PIR蛋白质序列数据库中注释的实验验证的翻译后修饰(PTM)。用于实体识别和本体开发的数据源包括蛋白质名称词典、单词标记词典、蛋白质名称标记的文献语料库沿着标记指南,以及基于PIRSF蛋白质家族名称的蛋白质本体。iProLINK可在http://pir.georgetown.edu/iprolink上免费访问,所有可下载文件都有超文本链接。(C)2004爱思唯尔有限公司保留所有权利。
The exponential growth of large-scale molecular sequence data and of the PubMed scientific literature has prompted active research in biological literature mining and information extraction to facilitate genome/proteome annotation and improve the quality of biological databases. Motivated by the promise of text mining methodologies, but at the same time, the lack of adequate curated data for training and benchmarking, the Protein Information Resource (PIR) has developed a resource for protein literature mining-iProLINK (integrated Protein Literature INformation and Knowledge). As PIR focuses its effort on the curation of the UniProt protein sequence database, the goal of iProLINK is to provide curated data sources that can be utilized for text mining research in the areas of bibliography mapping, annotation extraction. protein named entity recognition, and protein ontology development. The data sources for bibliography mapping and annotation extraction include mapped citations (PubMed ID to protein entry and feature line mapping) and annotation-tagged literature corpora. The latter includes several hundred abstracts and full-text articles tagged with experimentally validated post-translational modifications (PTMs) annotated in the PIR protein sequence database. The data sources for entity recognition and ontology development include a protein name dictionary, word token dictionaries, protein name-tagged literature corpora along with tagging guidelines, as well as a protein ontology based on PIRSF protein family names. iProLINK is freely accessible at http://pir.georgetown.edu/iprolink, with hypertext links for all downloadable files. (C) 2004 Elsevier Ltd. All rights reserved.