An adaptive classification model for peptide identification.

An adaptive classification model for peptide identification.
复制标题

DOI:
10.1186/1471-2164-16-s11-s1
复制
发表时间:
2015
期刊:
影响因子:
4.4
通讯作者:
Link A
Link A
中科院分区:
生物学2区
文献类型:
--
作者:
Liang X;Xia Z;Jian L;Niu X;Link A

文献摘要

被引文献

相似文献

肽序列分配是利用质谱/质谱技术鉴定蛋白质的中心任务。虽然已经开发了许多用于过滤目标肽谱匹配(psm)的数据库后搜索算法,但输出的psm之间的差异通常是显著的,留下了一些有争议的psm。目前的研究表明,仅使用判别函数很难将与诱饵弹相接近的目标弹相分离。在本文中,我们赋予每个目标PSM一个权值,表示其正确的可能性。我们采用基于支持向量机的学习模型来搜索每个目标PSM的最优权重,并开发了一个新的评分系统CRanker,对所有目标PSM进行排名。由于在日常数据库搜索中生成的大型PSM数据集,我们使用Cholesky分解技术来存储核矩阵以减少内存需求。与PeptideProphet和Percolator相比,CRanker在不同的数据集上发现了更多错误发现率相似的pms。CRanker在不同的测试集上表现出一致的性能,验证了所提模型的合理性。
Peptide sequence assignment is the central task in protein identification with MS/MS-based strategies. Although a number of post-database search algorithms for filtering target peptide spectrum matches (PSMs) have been developed, the discrepancy among the output PSMs is usually significant, remaining a few disputable PSMs. Current studies show that a number of target PSMs which are close to decoy PSMs can hardly be separated from those decoys by only using the discrimination function. In this paper, we assign each target PSM a weight showing its possibility of being correct. We employ a SVM-based learning model to search the optimal weight for each target PSM and develop a new score system, CRanker, to rank all target PSMs. Due to the large PSM datasets generated in routine database searches, we use the Cholesky factorization technique for storing a kernel matrix to reduce the memory requirement. Compared with PeptideProphet and Percolator, CRanker has identified more PSMs under similar false discover rates over different datasets. CRanker has shown consistent performance on different test sets, validated the reasonability the proposed model.