Scientific workflow optimization for improved peptide and protein identification.

Scientific workflow optimization for improved peptide and protein identification.
复制标题

优化科学工作流程,改进多肽和蛋白质鉴定。

DOI:
10.1186/s12859-015-0714-x
复制
发表时间:
2015-09-03
期刊:
影响因子:
3
通讯作者:
Palmblad M
Palmblad M
中科院分区:
生物学4区
文献类型:
--
作者:
Holl S;Mohammed Y;Zimmermann O;Palmblad M

文献摘要

被引文献

相似文献

肽谱匹配是基于质谱的蛋白质组学的大多数数据处理工作流程中的常见步骤。许多算法和软件包,无论是免费的还是商业的,已经开发出来解决这个问题。然而,这些算法通常需要用户选择仪器和样品相关参数,例如质量测量误差容限和错过的酶裂解次数。为了为特定数据集选择最佳算法和参数集,需要深入了解数据以及算法本身。因此,大多数研究人员倾向于使用默认参数,这些参数不一定是最佳的。我们已经为Taverna科学工作流程管理系统(http://ms-utils.org/Taverna_Optimization.pdf)应用了新的优化框架,以找到给定科学工作流程的最佳参数组合来执行肽谱匹配。优化本身是不平凡的,如通过在序列数据库搜索中允许较大质量测量误差时可以观察到的几种现象所证明的。科学工作流管理系统中嵌入的实时参数优化使专家和非专家能够从数据中提取最大量的信息。相同的工作流程可用于探索参数空间和比较算法,不仅用于肽谱匹配,还用于其他任务,例如保留时间预测。使用优化框架,我们能够了解如何获取数据以及探索的算法。我们观察到一种现象,确定许多氨损失b-离子光谱与N-末端焦谷氨酸和一个大的前体质量测量误差的肽。这些见解只能通过扩展优化框架探索的质量测量误差容限参数的共同范围来获得。
Peptide-spectrum matching is a common step in most data processing workflows for mass spectrometry-based proteomics. Many algorithms and software packages, both free and commercial, have been developed to address this task. However, these algorithms typically require the user to select instrument- and sample-dependent parameters, such as mass measurement error tolerances and number of missed enzymatic cleavages. In order to select the best algorithm and parameter set for a particular dataset, in-depth knowledge about the data as well as the algorithms themselves is needed. Most researchers therefore tend to use default parameters, which are not necessarily optimal. We have applied a new optimization framework for the Taverna scientific workflow management system (http://ms-utils.org/Taverna_Optimization.pdf) to find the best combination of parameters for a given scientific workflow to perform peptide-spectrum matching. The optimizations themselves are non-trivial, as demonstrated by several phenomena that can be observed when allowing for larger mass measurement errors in sequence database searches. On-the-fly parameter optimization embedded in scientific workflow management systems enables experts and non-experts alike to extract the maximum amount of information from the data. The same workflows could be used for exploring the parameter space and compare algorithms, not only for peptide-spectrum matching, but also for other tasks, such as retention time prediction. Using the optimization framework, we were able to learn about how the data was acquired as well as the explored algorithms. We observed a phenomenon identifying many ammonia-loss b-ion spectra as peptides with N-terminal pyroglutamate and a large precursor mass measurement error. These insights could only be gained with the extension of the common range for the mass measurement error tolerance parameters explored by the optimization framework.