Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasets.

Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasets.
复制标题

Hammock:一种隐藏的基于马尔可夫模型的肽聚类算法,以识别大型数据集中的蛋白质相互作用共识基序。

DOI:
10.1093/bioinformatics/btv522
复制
发表时间:
2016-01-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Muller P
Muller P
中科院分区:
其他
文献类型:
--
作者:
Krejci A;Hupp TR;Lexa M;Vojtesek B;Muller P

文献摘要

被引文献

相似文献

动机:蛋白质通常基于位于蛋白质表面无序区域的短线性基序识别其相互作用伙伴。研究这种基序的实验技术使用短肽来模拟相互作用蛋白质的结构特性。这些方法的持续发展允许大规模筛选,产生大量的肽序列,可能包含关于多种蛋白质-蛋白质相互作用的信息。这些数据集的处理对于研究蛋白质-蛋白质相互作用的大规模研究来说是一项复杂但必不可少的任务。结果如下:本文提出的软件工具能够快速识别来自各种来源的海量数据集中携带共享特异性基序的多个序列簇,并生成所识别簇的多个序列比对。该方法应用于先前发表的较小的数据集,包含不同类别的SH 3结构域的配体,以及一个新的,一个数量级的较大的数据集,包含几个单克隆抗体的表位。该软件成功地鉴定了模拟抗体靶标表位的序列簇,以及揭示抗体接受与原始表位序列的一些偏差的二级簇。另一个测试表明,处理更大的数据集在计算上是可行的。可用性和实施:Hammock是在GNU GPL v. 3许可下发布的,可以作为一个独立的程序(来自http://www.recamo.cz/en/software/hammock-cluster-peptides/)或作为Galaxy工具箱的工具(来自https://toolshed.g2.bx.psu.edu/view/hammock/hammock)免费获得。源代码可以从https://github.com/hammock-dev/hammock/releases下载。联系方式:muller@mou.cz补充信息:补充数据可从生物信息学在线网站获得。
Motivation: Proteins often recognize their interaction partners on the basis of short linear motifs located in disordered regions on proteins’ surface. Experimental techniques that study such motifs use short peptides to mimic the structural properties of interacting proteins. Continued development of these methods allows for large-scale screening, resulting in vast amounts of peptide sequences, potentially containing information on multiple protein-protein interactions. Processing of such datasets is a complex but essential task for large-scale studies investigating protein-protein interactions. Results: The software tool presented in this article is able to rapidly identify multiple clusters of sequences carrying shared specificity motifs in massive datasets from various sources and generate multiple sequence alignments of identified clusters. The method was applied on a previously published smaller dataset containing distinct classes of ligands for SH3 domains, as well as on a new, an order of magnitude larger dataset containing epitopes for several monoclonal antibodies. The software successfully identified clusters of sequences mimicking epitopes of antibody targets, as well as secondary clusters revealing that the antibodies accept some deviations from original epitope sequences. Another test indicates that processing of even much larger datasets is computationally feasible. Availability and implementation: Hammock is published under GNU GPL v. 3 license and is freely available as a standalone program (from http://www.recamo.cz/en/software/hammock-cluster-peptides/) or as a tool for the Galaxy toolbox (from https://toolshed.g2.bx.psu.edu/view/hammock/hammock). The source code can be downloaded from https://github.com/hammock-dev/hammock/releases. Contact: muller@mou.cz Supplementary information: Supplementary data are available at Bioinformatics online.