An iterative statistical approach to the identification of protein phosphorylation motifs from large-scale data sets

An iterative statistical approach to the identification of protein phosphorylation motifs from large-scale data sets
复制标题

DOI:
10.1038/nbt1146
复制
发表时间:
2005-11-01
影响因子:
46.9
通讯作者:
Gygi, SP
Gygi, SP
中科院分区:
工程技术1区
文献类型:
--
作者:
Schwartz, D;Gygi, SP

文献摘要

被引文献

相似文献

随着最近通过质谱法鉴定的蛋白质磷酸化位点呈指数增加,出现了了解这些位点周围的基序的独特机会。在这里,我们提出了一种算法,旨在从天然存在的磷酸化位点的大数据集中提取基序。该方法依赖于磷酸残基的内在比对以及通过与动态统计背景的迭代比较来提取基序。结果显示,从最近发表的丝氨酸、苏氨酸和酪氨酸磷酸化研究中鉴定出了数十种新颖和已知的磷酸化基序。当应用于语言数据集来测试该方法的多功能性时,该算法成功提取了数百个语言主题。该方法除了揭示已识别和尚未识别的激酶和模块化蛋白质结构域的共有序列之外,最终还可以用作确定感兴趣蛋白质中潜在磷酸化位点的工具。
With the recent exponential increase in protein phosphorylation sites identified by mass spectrometry, a unique opportunity has arisen to understand the motifs surrounding such sites. Here we present an algorithm designed to extract motifs from large data sets of naturally occurring phosphorylation sites. The methodology relies on the intrinsic alignment of phospho-residues and the extraction of motifs through iterative comparison to a dynamic statistical background. Results show the identification of dozens of novel and known phosphorylation motifs from recently published serine, threonine and tyrosine phosphorylation studies. When applied to a linguistic data set to test the versatility of the approach, the algorithm successfully extracted hundreds of language motifs. This method, in addition to shedding light on the consensus sequences of identified and as yet unidentified kinases and modular protein domains, may also eventually be used as a tool to determine potential phosphorylation sites in proteins of interest.