Detection and validation of non-synonymous coding SNPs from orthogonal analysis of shotgun proteomics data

Detection and validation of non-synonymous coding SNPs from orthogonal analysis of shotgun proteomics data
复制标题

DOI:
10.1021/pr0700908
复制
发表时间:
2007-01-01
影响因子:
4.4
通讯作者:
Stephenson, James L., Jr.
Stephenson, James L., Jr.
中科院分区:
生物学2区
文献类型:
--
作者:
Bunger, Maureen K.;Cargile, Benjamin J.;Stephenson, James L., Jr.

文献摘要

被引文献

相似文献

在现有的蛋白质组数据集的SNPs的结果的氨基酸取代的正交分析提供了一个重要的基础,新兴的领域,人口为基础的蛋白质组学。大规模的蛋白质组学数据集,来自鸟枪串联质谱分析复杂的细胞蛋白质混合物,包含许多未分配的光谱,可能对应于由SNP编码的替代等位基因。这项工作的目的是在LC-MS/MS鸟枪蛋白质组学数据集中识别可能代表编码非同义SNP(nsSNP)的串联MS谱。为此,我们从NCBI的dbSNP中发现的等位基因信息生成了胰蛋白酶肽数据库。我们用来自DU 4475乳腺肿瘤细胞的胰蛋白酶肽的串联MS光谱搜索了该数据库,所述胰蛋白酶肽在第一维中通过p/分级,在第二维中通过反相LC分级。我们总共鉴定了629个nsSNP,其中36个是在参考NCBI或IPI蛋白质数据库中未发现的替代SNP等位基因。SNP-肽的测序具有高的假阳性风险,这是由于修饰引起的质量偏移和由于基因组内相同肽的多种表示。在这项工作中,使用新的肽p/预测算法过滤假阳性,并使用通过随机取代类似大小的参考肽开发的诱饵数据库进行表征。通过相应基因组DNA测序的二次验证证实了10个SNP肽中的8个中存在预测的SNP。这项工作强调了将未分配的光谱解释为多态性的有用性高度依赖于检测和过滤假阳性的能力。
Orthogonal analysis of amino acid substitutions as a result of SNPs in existing proteomic datasets provides a critical foundation for the emerging field of population-based proteomics. Large-scale proteomics datasets, derived from shotgun tandem MS analysis of complex cellular protein mixtures, contain many unassigned spectra that may correspond to alternate alleles coded by SNPs. The purpose of this work was to identify tandem MS spectra in LC-MS/MS shotgun proteomics datasets that may represent coding nonsynonymous SNPs ( nsSNP). To this end, we generated a tryptic peptide database created from allelic information found in NCBI's dbSNP. We searched this database with tandem MS spectra of tryptic peptides from DU4475 breast tumor cells that had been fractioned by p/ in the first-dimension and reverse-phase LC in the second dimension. In all we identified 629 nsSNPs, of which 36 were of alternate SNP alleles not found in the reference NCBI or IPI protein databases. Searches for SNP-peptides carry a high risk of false positives due both to mass shifts caused by modifications and because of multiple representations of the same peptide within the genome. In this work, false positives were filtered using a novel peptide p/ prediction algorithm and characterized using a decoy database developed by random substitution of similarly sized reference peptides. Secondary validation by sequencing of corresponding genomic DNA confirmed the presence of the predicted SNP in 8 of 10 SNP-peptides. This work highlights that the usefulness of interpreting unassigned spectra as polymorphisms is highly reliant on the ability to detect and filter false positives.