The Protein Identifier Cross-Referencing (PICR) service: reconciling protein identifiers across multiple source databases.

The Protein Identifier Cross-Referencing (PICR) service: reconciling protein identifiers across multiple source databases.
复制标题

DOI:
10.1186/1471-2105-8-401
复制
发表时间:
2007-10-18
期刊:
影响因子:
3
通讯作者:
Hermjakob, Henning
Hermjakob, Henning
中科院分区:
生物学4区
文献类型:
--
作者:
Cote, Richard G;Jones, Philip;Martens, Lennart;Kerrien, Samuel;Reisinger, Florian;Lin, Quan;Leinonen, Rasko;Apweiler, Rolf;Hermjakob, Henning

文献摘要

被引文献

相似文献

每个主要的蛋白质数据库在分配蛋白质标识符时都有自己的惯例。解析那些指代相同蛋白质的各种可能不稳定的标识符是一项重大挑战。当试图统一由多个数据源的蛋白质标注的数据集,或者当源数据库使用一种蛋白质标识符而向数据提供者查询另一种时,这是一个常见问题。蛋白质标识符映射的部分解决方案是存在的,但它们局限于特定的物种或技术以及极少数的数据库。因此,我们没有找到一种通用性足够强且映射范围足够广以满足我们需求的解决方案。 我们创建了蛋白质标识符交叉引用(PICR)服务,这是一个网络应用程序,它为一种映射算法提供交互式和编程式(SOAP和REST)访问,该算法使用UniProt Archive(UniParc)作为数据仓库,基于与加载到UniParc中的来自70多个不同源数据库的蛋白质100%的序列同一性来提供蛋白质交叉引用。映射可以根据源数据库、分类学ID以及源数据库中的活性状态进行限制。用户可以复制/粘贴或上传包含蛋白质标识符或FASTA格式序列的文件,通过交互式界面获取映射。搜索结果可以在简单或详细的HTML表格中查看,或者作为逗号分隔值(CSV)或Microsoft Excel(XLS)文件下载,以便在本地数据库或电子表格中使用。或者,也可以使用SOAP接口将PICR功能集成到其他应用程序中,还有一个轻量级的REST接口。 我们提供一个公开可用的服务,它可以将蛋白质标识符和蛋白质序列交互式地映射到大多数常用的蛋白质数据库。可以通过符合标准的SOAP接口或轻量级的REST接口进行编程访问。PICR接口、文档和代码示例可在[具体网址未给出]获取。
Each major protein database uses its own conventions when assigning protein identifiers. Resolving the various, potentially unstable, identifiers that refer to identical proteins is a major challenge. This is a common problem when attempting to unify datasets that have been annotated with proteins from multiple data sources or querying data providers with one flavour of protein identifiers when the source database uses another. Partial solutions for protein identifier mapping exist but they are limited to specific species or techniques and to a very small number of databases. As a result, we have not found a solution that is generic enough and broad enough in mapping scope to suit our needs. We have created the Protein Identifier Cross-Reference (PICR) service, a web application that provides interactive and programmatic (SOAP and REST) access to a mapping algorithm that uses the UniProt Archive (UniParc) as a data warehouse to offer protein cross-references based on 100% sequence identity to proteins from over 70 distinct source databases loaded into UniParc. Mappings can be limited by source database, taxonomic ID and activity status in the source database. Users can copy/paste or upload files containing protein identifiers or sequences in FASTA format to obtain mappings using the interactive interface. Search results can be viewed in simple or detailed HTML tables or downloaded as comma-separated values (CSV) or Microsoft Excel (XLS) files suitable for use in a local database or a spreadsheet. Alternatively, a SOAP interface is available to integrate PICR functionality in other applications, as is a lightweight REST interface. We offer a publicly available service that can interactively map protein identifiers and protein sequences to the majority of commonly used protein databases. Programmatic access is available through a standards-compliant SOAP interface or a lightweight REST interface. The PICR interface, documentation and code examples are available at .