Cluster based prediction of PDZ-peptide interactions.

Cluster based prediction of PDZ-peptide interactions.
复制标题

DOI:
10.1186/1471-2164-15-s1-s5
复制
发表时间:
2014
期刊:
影响因子:
4.4
通讯作者:
Backofen R
Backofen R
中科院分区:
生物学2区
文献类型:
--
作者:
Kundu K;Backofen R

文献摘要

被引文献

相似文献

PDZ结构域是最混杂的蛋白质识别模块之一,其与短线性肽结合并在细胞信号传导中起重要作用。近年来,很少有高通量技术(如蛋白质微阵列筛选、噬菌体展示)被用于测定PDZ结构域的体外结合特异性。目前,许多计算方法可用于预测PDZ-肽相互作用,但它们通常提供结构域特异性模型和/或具有有限的结构域覆盖。在这里,我们组成了最大的一组来自人类,小鼠,苍蝇和蠕虫蛋白质组的PDZ结构域和定义的结合模型PDZ结构域家族,以提高域的覆盖率和预测特异性。为此,我们首先确定了一个新的138 PDZ家族,包括548 PDZ结构域从上述生物体,根据其序列同一性的有效聚类的基础上。对于43个PDZ家族,涵盖226个PDZ域与可用的交互数据,我们使用支持向量机方法建立了专门的模型。家族模型的优点是它们还可以用于确定与已知家族具有足够序列同一性的新表征的PDZ结构域的结合特异性。由于目前大多数实验方法只提供了积极的数据,我们必须处理类不平衡的问题。因此,为了丰富负类,我们引入了一种强大的半监督技术来生成高置信度的非交互数据。我们报告竞争力的预测性能相对于国家的最先进的方法。我们的方法有几个贡献。首先,我们表明,域覆盖率可以通过应用准确的聚类技术。其次,我们开发了一种基于半监督策略的方法来获得高置信度的阴性数据。第三,我们允许结合肽中氨基酸位置之间的高阶相关性。第四,我们的方法足够通用,并且很容易适用于其他肽识别模块,如SH2结构域,最后,我们对101个人类和102个小鼠PDZ结构域进行了全基因组预测,并发现了具有生物相关性的新型相互作用。我们将所有的预测模型和全基因组预测免费提供给科学界。本文的在线版本(doi:10.1186/1471 - 2164 - 15-S1-S5)包含补充材料,可供授权用户使用。
PDZ domains are one of the most promiscuous protein recognition modules that bind with short linear peptides and play an important role in cellular signaling. Recently, few high-throughput techniques (e.g. protein microarray screen, phage display) have been applied to determine in-vitro binding specificity of PDZ domains. Currently, many computational methods are available to predict PDZ-peptide interactions but they often provide domain specific models and/or have a limited domain coverage. Here, we composed the largest set of PDZ domains derived from human, mouse, fly and worm proteomes and defined binding models for PDZ domain families to improve the domain coverage and prediction specificity. For that purpose, we first identified a novel set of 138 PDZ families, comprising of 548 PDZ domains from aforementioned organisms, based on efficient clustering according to their sequence identity. For 43 PDZ families, covering 226 PDZ domains with available interaction data, we built specialized models using a support vector machine approach. The advantage of family-wise models is that they can also be used to determine the binding specificity of a newly characterized PDZ domain with sufficient sequence identity to the known families. Since most current experimental approaches provide only positive data, we have to cope with the class imbalance problem. Thus, to enrich the negative class, we introduced a powerful semi-supervised technique to generate high confidence non-interaction data. We report competitive predictive performance with respect to state-of-the-art approaches. Our approach has several contributions. First, we show that domain coverage can be increased by applying accurate clustering technique. Second, we developed an approach based on a semi-supervised strategy to get high confidence negative data. Third, we allowed high order correlations between the amino acid positions in the binding peptides. Fourth, our method is general enough and will easily be applicable to other peptide recognition modules such as SH2 domains and finally, we performed a genome-wide prediction for 101 human and 102 mouse PDZ domains and uncovered novel interactions with biological relevance. We make all the predictive models and genome-wide predictions freely available to the scientific community. The online version of this article (doi:10.1186/1471-2164-15-S1-S5) contains supplementary material, which is available to authorized users.