Using high-abundance proteins as guides for fast and effective peptide/protein identification from human gut metaproteomic data.

Using high-abundance proteins as guides for fast and effective peptide/protein identification from human gut metaproteomic data.
复制标题

DOI:
10.1186/s40168-021-01035-8
复制
发表时间:
2021-04-01
期刊:
影响因子:
15.5
通讯作者:
Ye Y
Ye Y
中科院分区:
生物学1区
文献类型:
--
作者:
Stamboulian M;Li S;Ye Y

文献摘要

参考文献

相似文献

最近的一些重大努力显着扩大了人类相关细菌基因组的收集,现在包含数千个实体,包括参考完整/草稿基因组和宏基因组组装基因组(MAG)。这些基因组为研究人类相关微生物组的功能及其与人类健康和疾病的关系提供了有用的资源。这些基因组的应用之一是在无法获得匹配的宏基因组/宏转录组数据时,为宏蛋白质组研究中的数据库搜索提供通用参考。然而,更多的参考基因组集合可能不一定会导致更好的肽/蛋白质识别,因为搜索空间的增加通常会导致谱图-肽匹配更少,更不用说计算时间的急剧增加。视频摘要 在这里,我们提出了一种新方法,该方法使用两个步骤来优化参考基因组和 MAG 的使用,作为人类肠道宏蛋白质组 MS/MS 数据分析的通用参考。第一步是仅使用高丰度蛋白 (HAP)(即核糖体蛋白和延伸因子)进行宏蛋白质组 MS/MS 数据库搜索,并根据鉴定结果推导出潜在微生物群落的分类组成。第二步是通过包含来自已识别的丰富物种的所有蛋白质来扩展搜索数据库。我们将我们的方法称为 HAPiID(HAP 引导的宏蛋白质组学鉴定)。我们使用先前研究中的人类肠道宏蛋白质组数据集测试了我们的方法,并将其与最先进的参考数据库搜索方法 MetaPro-IQ 进行比较,用于研究人类肠道微生物群的宏蛋白质组鉴定。我们的结果表明,我们的两步方法不仅执行速度明显更快,而且能够识别更多的肽。我们进一步证明了 HAPiID 在揭示单个人类相关细菌物种(一次一个或几个物种)的蛋白质谱方面的应用,使用宏蛋白质组数据。 HAP 引导分析方法为构建宏蛋白质组数据分析的目标数据库提供了一种新颖有效的方法。基于这种方法构建的 HAPiID 管道为分析人类肠道相关的宏蛋白质组数据提供了通用工具。在线版本包含可在 (10.1186/s40168-021-01035-8) 获取的补充材料。
A few recent large efforts significantly expanded the collection of human-associated bacterial genomes, which now contains thousands of entities including reference complete/draft genomes and metagenome assembled genomes (MAGs). These genomes provide useful resource for studying the functionality of the human-associated microbiome and their relationship with human health and diseases. One application of these genomes is to provide a universal reference for database search in metaproteomic studies, when matched metagenomic/metatranscriptomic data are unavailable. However, a greater collection of reference genomes may not necessarily result in better peptide/protein identification because the increase of search space often leads to fewer spectrum-peptide matches, not to mention the drastic increase of computation time. Video Abstract Here, we present a new approach that uses two steps to optimize the use of the reference genomes and MAGs as the universal reference for human gut metaproteomic MS/MS data analysis. The first step is to use only the high-abundance proteins (HAPs) (i.e., ribosomal proteins and elongation factors) for metaproteomic MS/MS database search and, based on the identification results, to derive the taxonomic composition of the underlying microbial community. The second step is to expand the search database by including all proteins from identified abundant species. We call our approach HAPiID (HAPs guided metaproteomics IDentification). We tested our approach using human gut metaproteomic datasets from a previous study and compared it to the state-of-the-art reference database search method MetaPro-IQ for metaproteomic identification in studying human gut microbiota. Our results show that our two-steps method not only performed significantly faster but also was able to identify more peptides. We further demonstrated the application of HAPiID to revealing protein profiles of individual human-associated bacterial species, one or a few species at a time, using metaproteomic data. The HAP guided profiling approach presents a novel effective way for constructing target database for metaproteomic data analysis. The HAPiID pipeline built upon this approach provides a universal tool for analyzing human gut-associated metaproteomic data. The online version contains supplementary material available at (10.1186/s40168-021-01035-8).
DOI: 10.1016/j.cels.2018.08.009
发表时间: 2018-10-24
期刊: Cell systems
影响因子: 9.3
作者:
Beyter D;Lin MS;Yu Y;Pieper R;Bafna V
通讯作者: Bafna V
DOI: 10.1038/ismej.2011.159
发表时间: 2012-05
期刊: The ISME journal
影响因子: --
作者:
通讯作者: --
DOI: 10.1371/journal.pgen.1000556
发表时间: 2009-07
期刊: PLoS genetics
影响因子: 4.5
作者:
Hershberg R;Petrov DA
通讯作者: Petrov DA
DOI: 10.1093/nar/gkv397
发表时间: 2015-07-01
影响因子: 14.9
作者:
Finn RD;Clements J;Arndt W;Miller BL;Wheeler TJ;Schreiber F;Bateman A;Eddy SR
通讯作者: Eddy SR
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W