Protein and peptide identification algorithms using MS for use in high‐throughput, automated pipelines

Protein and peptide identification algorithms using MS for use in high‐throughput, automated pipelines
复制标题

DOI:
10.1002/pmic.200402091
复制
发表时间:
2005-11
期刊:
影响因子:
3.4
通讯作者:
I. Shadforth;Daniel J. Crowther;C. Bessant
I. Shadforth;Daniel J. Crowther;C. Bessant
中科院分区:
生物学3区
文献类型:
--
作者:
I. Shadforth;Daniel J. Crowther;C. Bessant

文献摘要

被引文献

相似文献

目前的蛋白质组学实验可以非常快地产生大量数据,但数据分析能力无法与之匹敌。虽然最近已经有一些关于利用MS鉴定多肽和蛋白质的方法的不同方面的综述,但关于哪种方法最适合于他们所提出的任务或在他们所提议的任务中最有效的比较是不容易获得的。随着对高通量、自动化的多肽和蛋白质鉴定系统的需求增加,这类管道的创建者需要能够选择在准确性和计算效率方面都表现良好的算法。因此,本文对目前可用的PMF核心算法、利用MS/MS进行数据库搜索、序列标签搜索和从头测序进行了综述。我们还评估了其中一些算法的相对性能。由于文献中对这类信息的报道有限,我们得出结论,有必要采用一种基于免费可用的数据集的关于新的肽和蛋白质鉴定算法的性能的标准化报告系统。我们继续就这些数据集的格式和内容提出我们的初步建议。
Current proteomics experiments can generate vast quantities of data very quickly, but this has not been matched by data analysis capabilities. Although there have been a number of recent reviews covering various aspects of peptide and protein identification methods using MS, comparisons of which methods are either the most appropriate for, or the most effective at, their proposed tasks are not readily available. As the need for high‐throughput, automated peptide and protein identification systems increases, the creators of such pipelines need to be able to choose algorithms that are going to perform well both in terms of accuracy and computational efficiency. This article therefore provides a review of the currently available core algorithms for PMF, database searching using MS/MS, sequence tag searches and de novo sequencing. We also assess the relative performances of a number of these algorithms. As there is limited reporting of such information in the literature, we conclude that there is a need for the adoption of a system of standardised reporting on the performance of new peptide and protein identification algorithms, based upon freely available datasets. We go on to present our initial suggestions for the format and content of these datasets.