High-throughput identification of proteins and unanticipated sequence modifications using a mass-based alignment algorithm for MS/MS de novo sequencing results

High-throughput identification of proteins and unanticipated sequence modifications using a mass-based alignment algorithm for MS/MS de novo sequencing results
复制标题

DOI:
10.1021/ac035258x
复制
发表时间:
2004-04-15
影响因子:
7.4
通讯作者:
Nagalla, SR
Nagalla, SR
中科院分区:
化学1区
文献类型:
--
作者:
Searle, BC;Dasari, S;Nagalla, SR

文献摘要

被引文献

相似文献

随着越来越多用于解释高质量精度串联质谱(MS/MS)数据的从头测序算法的可用性,越来越需要从从头测序结果中准确识别蛋白质的程序。从肽的串联质谱中衍生的从头序列通常包含模棱两可的区域,其中无法确定确切的氨基酸顺序。这给序列比对算法带来的一个问题是难以区分由于从头测序错误与实际基因组序列变异和翻译后修饰引起的差异。我们提出了一种新颖的、基于质量的序列比对方法,作为一个名为OpenSea的程序来实现,以解决这些问题。在这种方法中,de novo和数据库序列被解释为残基的质量,并且质量而不是氨基酸编码进行比较。为了提供进一步的灵活性,可以将质量分组排列,这可以解决许多从头测序错误。OpenSea的性能用三种类型的数据进行测试:已知蛋白质的混合物,通常含有序列变异的未知蛋白质的混合物,以及翻译后修饰的已知蛋白质的混合物。在这三种情况下,我们证明了OpenSea可以比常用的数据库搜索程序(SEQUEST和ProteinLynx)识别更多的肽和蛋白质,同时在高通量环境中准确定位序列变异位点和意想不到的翻译后修饰。
With the increasing availability of de novo sequencing algorithms for interpreting high-mass accuracy tandem mass spectrometty (MS/MS) data, there is a growing need for programs that accurately identify proteins from de novo sequencing results. De novo sequences derived from tandem mass spectra of peptides often contain ambiguous regions where the exact amino acid order cannot be determined. One problem this poses for sequence alignment algorithms is the difficulty in distinguishing discrepancies due to de novo sequencing errors from actual genomic sequence variation and posttranslational modifications. We present a novel, mass-based approach to sequence alignment, implemented as a program called OpenSea, to resolve these problems. In this approach, de novo and database sequences are interpreted as masses of residues, and the masses, rather than the amino acid codes, are compared. To provide further flexibility, the masses can be aligned in groups, which can resolve many de novo sequencing errors. The performance of OpenSea was tested with three types of data: a mixture of known proteins, a mixture of unknown proteins that commonly contain sequence variations, and a mixture of posttranslationally modified known proteins. In all three cases, we demonstrate that OpenSea can identify more peptides and proteins than commonly used database-searching programs (SEQUEST and ProteinLynx) while accurately locating sequence variation sites and unanticipated posttranslational modifications in a high-throughput environment.