Implications of strain- and species-level sequence divergence for community and isolate shotgun proteomic analysis

Implications of strain- and species-level sequence divergence for community and isolate shotgun proteomic analysis
复制标题

DOI:
10.1021/pr0701005
复制
发表时间:
2007-08-01
影响因子:
4.4
通讯作者:
Banfield, Jillian F.
Banfield, Jillian F.
中科院分区:
生物学2区
文献类型:
--
作者:
Denef, Vincent J.;Shah, Manesh B.;Banfield, Jillian F.

文献摘要

被引文献

相似文献

近年来微生物基因组测序的兴起,以及基于高通量液相色谱-质谱(LC/LC-MS/MS)的蛋白质组学的发展,提出了一个问题,即一个菌株或环境样品的基因组信息可以在多大程度上用于分析相关菌株或样品的蛋白质组学。即使随着测序成本的降低,获得每个菌株或样本的基因组序列仍然是不切实际的。在这里,我们使用基于概率的模型和受实验数据约束的随机突变模拟模型来评估样本和基因组数据库之间的氨基酸差异如何影响散弹枪蛋白质组学。为了评估突变非随机分布的影响,我们还利用测序分离株的硅肽数据评估了鉴定水平,这些分离株的平均氨基酸身份(AAI)在76 - 98%之间变化。我们将预测结果与一个样本的实验蛋白质鉴定水平进行了比较,该样本使用一个数据库进行评估,该数据库包括优势生物体和密切相关变体的基因组信息(95% AAI)。模型的范围设定了蛋白质组学实验中一半的蛋白质在样品和数据库中同源物之间的AAI为77-92%的界限。与这一预测一致,实验数据表明,在90% AAI时,可识别的蛋白质损失了一半。进一步的分析表明,每1%氨基酸分化,初始蛋白质覆盖率降低6.4%,总鉴定损失为86% AAI。因此,霰弹枪蛋白质组学能够跨菌株鉴定,但避免了大多数跨物种的假阳性。
The recent surge in microbial genomic sequencing, combined with the development of high-throughput liquid chromatography-mass-spectrometry-based (LC/LC-MS/MS) proteomics, has raised the question of the extent to which genomic information of one strain or environmental sample can be used to profile proteomes of related strains or samples. Even with decreasing sequencing costs, it remains impractical to obtain genomic sequence for every strain or sample analyzed. Here, we evaluate how shotgun proteomics is affected by amino acid divergence between the sample and the genomic database using a probability-based model and a random mutation simulation model constrained by experimental data. To assess the effects of nonrandom distribution of mutations, we also evaluated identification levels using in silico peptide data from sequenced isolates with average amino acid identities (AAI) varying between 76 and 98%. We compared the predictions to experimental protein identification levels for a sample that was evaluated using a database that included genomic information for the dominant organism and for a closely related variant (95% AAI). The range of models set the boundaries at which half of the proteins in a proteomic experiment can be identified to be 77-92% AAI between orthologs in the sample and database. Consistent with this prediction, experimental data indicated loss of half the identifiable proteins at 90% AAI. Additional analysis indicated a 6.4% reduction of the initial protein coverage per 1% amino acid divergence and total identification loss at 86% AAI. Consequently, shotgun proteomics is capable of cross-strain identifications but avoids most cross-species false positives.