Identifying HLA supertypes by learning distance functions

Identifying HLA supertypes by learning distance functions
复制标题

DOI:
10.1093/bioinformatics/btl324
复制
发表时间:
2007-01-15
期刊:
影响因子:
5.8
通讯作者:
Yanover, Chen
Yanover, Chen
中科院分区:
生物学3区
文献类型:
--
作者:
Hertz, Tomer;Yanover, Chen

文献摘要

被引文献

相似文献

动机:基于表位的疫苗的开发主要依赖于将人类白细胞抗原(HLA)分子分类为具有相似肽结合特异性的组的能力,称为超型。在他们的开创性工作中,Sette和Sidney定义了9个HLA I类超型,并声称这些超型几乎完美地覆盖了HLA I类分子的全部功能。HLA等位基因具有高度多态性和多基因性,因此通过实验对这些分子进行超型分类目前是一项不可能完成的任务。最近,针对这一任务提出了许多计算方法。这些方法是基于定义蛋白质相似性测量,从结合肽的分析或从蛋白质本身的分析得出。结果:在本文中,我们定义了基于学习距离函数的肽衍生和蛋白质衍生相似性度量。肽衍生的测量是使用肽-肽距离函数来定义的,该函数是使用已知结合和非结合肽的信息来学习的。蛋白质衍生的相似性度量是使用蛋白质-蛋白质距离函数来定义的,该函数是使用先前由Sette和Sidney(1999)分类为超型的等位基因信息来学习的。我们将这两种互补方法得到的分类与先前建议的分类方法进行比较。总的来说,我们的结果与Sette和Sidney(1999)提出的分类以及Buus等人(2004)报告的分类非常一致。我们提出的基于距离的方法的主要优势是它利用了两种不同的重要免疫学信息来源——hla等位基因和已知与这些等位基因结合或不结合的肽。由于我们的每一种距离测量方法都是使用不同的信息源进行训练的,因此它们的组合可以为超型等位基因提供更可靠的分类。
Motivation: The development of epitope-based vaccines crucially relies on the ability to classify Human Leukocyte Antigen (HLA) molecules into sets that have similar peptide binding specificities, termed supertypes. In their seminal work, Sette and Sidney defined nine HLA class I supertypes and claimed that these provide an almost perfect coverage of the entire repertoire of HLA class I molecules.HLA alleles are highly polymorphic and polygenic and therefore experimentally classifying each of these molecules to supertypes is at present an impossible task. Recently, a number of computational methods have been proposed for this task. These methods are based on defining protein similarity measures, derived from analysis of binding peptides or from analysis of the proteins themselves.Results: In this paper we define both peptide derived and protein derived similarity measures, which are based on learning distance functions. The peptide derived measure is defined using a peptide-peptide distance function, which is learned using information about known binding and non-binding peptides. The protein derived similarity measure is defined using a protein-protein distance function, which is learned using information about alleles previously classified to supertypes by Sette and Sidney (1999). We compare the classification obtained by these two complimentary methods to previously suggested classification methods. In general, our results are in excellent agreement with the classifications suggested by Sette and Sidney (1999) and with those reported by Buus et al. (2004).The main important advantage of our proposed distance-based approach is that it makes use of two different and important immunological sources of information-HLA alleles and peptides that are known to bind or not bind to these alleles. Since each of our distance measures is trained using a different source of information, their combination can provide a more confident classification of alleles to supertypes.