Interpretation of mass spectrometry data for high-throughput proteomics

Interpretation of mass spectrometry data for high-throughput proteomics
复制标题

DOI:
10.1007/s00216-003-1995-x
复制
发表时间:
2003-08-01
影响因子:
4.3
通讯作者:
Blueggel, M
Blueggel, M
中科院分区:
化学2区
文献类型:
--
作者:
Chamrad, DC;Koerting, G;Blueggel, M

文献摘要

被引文献

相似文献

蛋白质组学的最新发展揭示了生物信息学的瓶颈:获得的质谱数据的高质量解释。每天生成数千个质谱的能力,以及对这种能力的需求,使得手工方法不足以进行分析,并强调需要将专家用户的高级能力转移到复杂的质谱解释算法中。在目前的高通量蛋白质组学研究中,鉴定率不仅仅是仪器的问题。我们提出了用于高通量PMF鉴定的软件,它能够以更高的速率进行稳健和自信的蛋白质鉴定。这是通过自动校准、峰值抑制和使用元搜索方法实现的,该方法采用了各种PMF搜索引擎。自动定标是一种动态的、依赖光谱信息的算法,它结合了各种已知的定标方法,迭代地建立一个最优定标。峰值拒绝算法通过使用自动生成的和数据集相关的排除列表来过滤与所分析蛋白质无关的信号。在“元搜索”中,几个已知的PMF搜索引擎被触发,它们的结果通过使用元分数被合并。meta评分的显著性是通过模拟PMF识别的10,000个人工光谱来评估的,这些光谱类似于接近测量数据集的数据情况。通过这种模拟,元分数作为一种统计度量与期望值相关联。该软件是蛋白质组数据库ProteinScape的一部分,该数据库将MS数据衍生的信息链接到其他相关的蛋白质组学数据。我们用1891 PMF光谱的质谱数据证明了该系统的性能。由于自动校准和峰值抑制,识别率从6%增加到44%。
Recent developments in proteomics have revealed a bottleneck in bioinformatics: high-quality interpretation of acquired MS data. The ability to generate thousands of MS spectra per day, and the demand for this, makes manual methods inadequate for analysis and underlines the need to transfer the advanced capabilities of an expert human user into sophisticated MS interpretation algorithms. The identification rate in current high-throughput proteomics studies is not only a matter of instrumentation. We present software for high-throughput PMF identification, which enables robust and confident protein identification at higher rates. This has been achieved by automated calibration, peak rejection, and use of a meta search approach which employs various PMF search engines. The automatic calibration consists of a dynamic, spectral information-dependent algorithm, which combines various known calibration methods and iteratively establishes an optimised calibration. The peak rejection algorithm filters signals that are unrelated to the analysed protein by use of automatically generated and dataset-dependent exclusion lists. In the "meta search" several known PMF search engines are triggered and their results are merged by use of a meta score. The significance of the meta score was assessed by simulation of PMF identification with 10,000 artificial spectra resembling a data situation close to the measured dataset. By means of this simulation the meta score is linked to expectation values as a statistical measure. The presented software is part of the proteome database ProteinScape which links the information derived from MS data to other relevant proteomics data. We demonstrate the performance of the presented system with MS data from 1891 PMF spectra. As a result of automatic calibration and peak rejection the identification rate increased from 6% to 44%.