MoMo: discovery of statistically significant post-translational modification motifs

MoMo: discovery of statistically significant post-translational modification motifs
复制标题

DOI:
10.1093/bioinformatics/bty1058
复制
发表时间:
2019-08-15
期刊:
影响因子:
5.8
通讯作者:
Bailey, Timothy L.
Bailey, Timothy L.
中科院分区:
生物学3区
文献类型:
--
作者:
Cheng, Alice;Grant, Charles E.;Bailey, Timothy L.

文献摘要

被引文献

相似文献

动机:蛋白质的翻译后修饰(PTM)与许多重要的生物学功能相关,并且可以使用串联质谱进行高通量鉴定。许多PTM与称为“基序”的短序列模式相关,这些模式有助于定位修饰酶。因此,已经设计了许多算法来从质谱数据中识别这些基序。准确的统计置信度估计发现的图案是至关重要的正确的解释和下游的实验validation.Results的设计:我们描述了一种方法分配统计置信度估计PTM图案,我们证明,这种方法提供了准确的P值模拟和真实的数据。我们的方法是在莫莫中实现的,MoMo是一种软件工具,用于在我们作为Web服务器和可下载的源代码提供的PTM集合中发现图案。莫莫重新实现了两个最广泛使用的PTM基序发现算法-motif-x和MoDL-同时提供了许多增强功能。相对于基序-x,莫莫提供了改进的统计置信度估计和更准确的基序得分计算。莫莫网络服务器提供更多的蛋白质组数据库,更多的输入格式,更大的输入和更长的运行时间比motif-x网络服务器。最后,我们的研究表明,由基序-x产生的置信度估计是不准确的。这种不准确性部分源于从未改组的蛋白质组数据库中提取“背景”肽的常见做法。因此,我们的研究结果表明,许多使用motif-x来寻找motifs的论文可能报告了缺乏统计支持的结果。
Motivation: Post-translational modifications (PTMs) of proteins are associated with many significant biological functions and can be identified in high throughput using tandem mass spectrometry. Many PTMs are associated with short sequence patterns called 'motifs' that help localize the modifying enzyme. Accordingly, many algorithms have been designed to identify these motifs from mass spectrometry data. Accurate statistical confidence estimates for discovered motifs are critically important for proper interpretation and in the design of downstream experimental validation.Results: We describe a method for assigning statistical confidence estimates to PTM motifs, and we demonstrate that this method provides accurate P-values on both simulated and real data. Our methods are implemented in MoMo, a software tool for discovering motifs among sets of PTMs that we make available as a web server and as downloadable source code. MoMo re-implements the two most widely used PTM motif discovery algorithms-motif-x and MoDL-while offering many enhancements. Relative to motif-x, MoMo offers improved statistical confidence estimates and more accurate calculation of motif scores. The MoMo web server offers more proteome databases, more input formats, larger inputs and longer running times than the motif-x web server. Finally, our study demonstrates that the confidence estimates produced by motif-x are inaccurate. This inaccuracy stems in part from the common practice of drawing 'background' peptides from an unshuffled proteome database. Our results thus suggest that many of the papers that use motif-x to find motifs may be reporting results that lack statistical support.