A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry data

A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry data
复制标题

DOI:
10.1186/1471-2105-7-222
复制
发表时间:
2006-04-26
期刊:
影响因子:
3
通讯作者:
Chen, Runsheng
Chen, Runsheng
中科院分区:
生物学4区
文献类型:
--
作者:
Zhang, Zhuo;Sun, Shiwei;Chen, Runsheng

文献摘要

被引文献

相似文献

背景:串联质谱(MS/ MS)是蛋白质鉴定的有力工具。结果:提出了一种新的评分函数,沿着了评价函数性能置信度的准则,并在此基础上提出了一种新的评分函数。通过学习子离子的类型和产生它们的概率,为每个候选肽生成一个假设的光谱。然后引入相对熵来度量假设光谱与观测光谱之间的相似性。基于极值分布(EVD)理论,选择阈值来区分真实肽分配与随机肽分配。公共MS/ MS数据集上的测试表明,该方法比著名的SEQUEST.Conclusion:一个可靠的识别蛋白质的光谱承诺串联质谱更有效地应用于蛋白质组具有高的复杂性。
Background: Tandem mass spectrometry ( MS/ MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra.Results: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution ( EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/ MS dataset demonstrated that this method performs better than the well-known SEQUEST.Conclusion: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity.