Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach

Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach
复制标题

DOI:
10.1093/bioinformatics/bth100
复制
发表时间:
2004-06-12
期刊:
影响因子:
5.8
通讯作者:
Lund, O
Lund, O
中科院分区:
生物学3区
文献类型:
--
作者:
Nielsen, M;Lundegaard, C;Lund, O

文献摘要

被引文献

相似文献

动机:预测哪些肽将结合特定的主要组织相容性复合物 (MHC),是识别适合作为候选疫苗的潜在 T 细胞表位的重要一步。 MHC II 类结合肽具有广泛的长度分布,使此类预测变得复杂。因此,识别正确的比对是识别 MHC II 类结合基序核心的关键部分。在这种情况下,我们希望描述一种新颖的吉布斯基序采样器方法,非常适合识别此类弱序列基序。该方法基于吉布斯采样方法,并结合了针对识别 MHC I 类和 II 类结合基序的任务而优化的新功能。该方法在一组序列中定位结合基序,并根据权重矩阵表征该基序。随后,权重矩阵可用于有效识别潜在的MHC结合肽并指导合理的疫苗设计过程。结果:我们将基序采样器方法应用于MHC II类结合的复杂问题。该方法的输入是从 SYFPEITHI 和 MHCPEP 公共数据库中提取的氨基酸肽序列,已知其与 MHC II 类复合物 HLA-DR4(B1*0401) 结合。预先识别结合基序中信息丰富的(锚定)位置可以提高吉布斯采样器的预测性能。类似地,从次优解决方案的整体平均值获得的一致解决方案被证明优于使用单一最优解决方案。在大规模基准计算中,使用相对操作特征曲线 (ROC) 图对性能进行量化,并且我们将性能与 TEPITOPE 方法和使用 ClustalW 传统对齐算法导出的权重矩阵的性能进行了详细比较。计算表明,Gibbs采样器的预测性能高于ClustalW,并且在大多数情况下也高于TEPITOPE方法。
Motivation: Prediction of which peptides will bind a specific major histocompatibility complex (MHC) constitutes an important step in identifying potential T-cell epitopes suitable as vaccine candidates. MHC class II binding peptides have a broad length distribution complicating such predictions. Thus, identifying the correct alignment is a crucial part of identifying the core of an MHC class II binding motif. In this context, we wish to describe a novel Gibbs motif sampler method ideally suited for recognizing such weak sequence motifs. The method is based on the Gibbs sampling method, and it incorporates novel features optimized for the task of recognizing the binding motif of MHC classes I and II. The method locates the binding motif in a set of sequences and characterizes the motif in terms of a weight-matrix. Subsequently, the weight-matrix can be applied to identifying effectively potential MHC binding peptides and to guiding the process of rational vaccine design.Results: We apply the motif sampler method to the complex problem of MHC class II binding. The input to the method is amino acid peptide sequences extracted from the public databases of SYFPEITHI and MHCPEP and known to bind to the MHC class II complex HLA-DR4(B1*0401). Prior identification of information-rich (anchor) positions in the binding motif is shown to improve the predictive performance of the Gibbs sampler. Similarly, a consensus solution obtained from an ensemble average over suboptimal solutions is shown to outperform the use of a single optimal solution. In a large-scale benchmark calculation, the performance is quantified using relative operating characteristics curve (ROC) plots and we make a detailed comparison of the performance with that of both the TEPITOPE method and a weight-matrix derived using the conventional alignment algorithm of ClustalW. The calculation demonstrates that the predictive performance of the Gibbs sampler is higher than that of ClustalW and in most cases also higher than that of the TEPITOPE method.