Simultaneous alignment and clustering of peptide data using a Gibbs sampling approach

Simultaneous alignment and clustering of peptide data using a Gibbs sampling approach
复制标题

DOI:
10.1093/bioinformatics/bts621
复制
发表时间:
2013-01-01
期刊:
影响因子:
5.8
通讯作者:
Nielsen, Morten
Nielsen, Morten
中科院分区:
生物学3区
文献类型:
--
作者:
Andreatta, Massimo;Lund, Ole;Nielsen, Morten

文献摘要

被引文献

相似文献

动机:识别短肽片段的蛋白质在细胞信号传导中起核心作用。由于高通量技术,可以使用大型肽库以显着更低的成本和时间研究肽结合蛋白的特异性。然而,这样大的肽数据集的解释,是一个复杂的任务,特别是当数据包含多个受体结合基序,和/或基序被发现在不同的位置内distinct peptide.Results:本文提出的算法,吉布斯采样的基础上,确定多个特异性肽数据同时执行两个基本任务:肽数据的对齐和聚类。我们将该方法应用于具有不同复杂程度的肽数据集面板中的去卷积结合基序,从最简单的预对齐固定长度肽的情况到可变长度的未对齐肽数据集的情况。本文中描述的示例应用包括针对不同MHC I类和II类等位基因的结合剂的混合物、针对SH 3结构域的不同类别的配体以及HLA-A*02:01分子的亚特异性。
Motivation: Proteins recognizing short peptide fragments play a central role in cellular signaling. As a result of high-throughput technologies, peptide-binding protein specificities can be studied using large peptide libraries at dramatically lower cost and time. Interpretation of such large peptide datasets, however, is a complex task, especially when the data contain multiple receptor binding motifs, and/or the motifs are found at different locations within distinct peptides.Results: The algorithm presented in this article, based on Gibbs sampling, identifies multiple specificities in peptide data by performing two essential tasks simultaneously: alignment and clustering of peptide data. We apply the method to de-convolute binding motifs in a panel of peptide datasets with different degrees of complexity spanning from the simplest case of pre-aligned fixed-length peptides to cases of unaligned peptide datasets of variable length. Example applications described in this article include mixtures of binders to different MHC class I and class II alleles, distinct classes of ligands for SH3 domains and sub-specificities of the HLA-A*02:01 molecule.