Analysis of the tryptic search space in UniProt databases

Analysis of the tryptic search space in UniProt databases
复制标题

DOI:
10.1002/pmic.201400227
复制
发表时间:
2015-01-01
期刊:
影响因子:
3.4
通讯作者:
Martin, Maria J.
Martin, Maria J.
中科院分区:
生物学3区
文献类型:
--
作者:
Alpi, Emanuele;Griss, Johannes;Martin, Maria J.

文献摘要

被引文献

相似文献

在这篇文章中,我们提供了一个全面的研究内容的通用蛋白质资源(UniProt)蛋白质数据集的人类和小鼠。将UniProtKB (UniProt知识库)完整蛋白质组集的色氨酸搜索空间与来自UniProtKB的其他数据集以及相应的国际蛋白质索引、参考序列、Ensembl和UniRef100(其中UniRef是UniProt参考簇)生物特异性数据集进行比较。本研究评估了UniProtKB中注释的所有蛋白形式(包括规范序列和同种异构体)。此外,在UniProtKB中标注的天然和疾病相关氨基酸变异也被纳入评估。对每个数据集的肽单一性也进行了评估。此外,还将UniProtKB数据集中的肽信息与主要MS-based蛋白质组学知识库中可用的肽水平鉴定进行了比较。鉴定这些库中观察到的肽是蛋白质数据库的重要信息来源,因为它们为其他预测的蛋白质的存在提供了支持证据。同样,知识库可以使用UniProtKB中提供的信息来指导对感兴趣的特定肽/蛋白质集的再处理工作。总之,我们提供了关于UniProt提供的不同生物特异性序列数据集的全面信息,以及每个数据集的优缺点,以及基于ms的自下而上蛋白质组学工作流程的搜索空间。分析的目的是为UniProt和其他蛋白质数据库的色氨酸搜索空间提供一个清晰的视图,使科学家能够选择最适合他们目的的那些。
In this article, we provide a comprehensive study of the content of the Universal Protein Resource (UniProt) protein data sets for human and mouse. The tryptic search spaces of the UniProtKB (UniProt knowledgebase) complete proteome sets were compared with other data sets from UniProtKB and with the corresponding International Protein Index, reference sequence, Ensembl, and UniRef100 (where UniRef is UniProt reference clusters) organism-specific data sets. All protein forms annotated in UniProtKB (both the canonical sequences and isoforms) were evaluated in this study. In addition, natural and disease-associated amino acid variants annotated in UniProtKB were included in the evaluation. The peptide unicity was also evaluated for each data set. Furthermore, the peptide information in the UniProtKB data sets was also compared against the available peptide-level identifications in the main MS-based proteomics repositories. Identifying the peptides observed in these repositories is an important resource of information for protein databases as they provide supporting evidence for the existence of otherwise predicted proteins. Likewise, the repositories could use the information available in UniProtKB to direct reprocessing efforts on specific sets of peptides/proteins of interest. In summary, we provide comprehensive information about the different organism-specific sequence data sets available from UniProt, together with the pros and cons for each, in terms of search space for MS-based bottom-up proteomics workflows. The aim of the analysis is to provide a clear view of the tryptic search space of UniProt and other protein databases to enable scientists to select those most appropriate for their purposes.