A community resource benchmarking predictions of peptide binding to MHC-I molecules.

A community resource benchmarking predictions of peptide binding to MHC-I molecules.
复制标题

肽与MHC-I分子结合的社区资源基准预测。

DOI:
10.1371/journal.pcbi.0020065
复制
发表时间:
2006-06-09
影响因子:
4.3
通讯作者:
Sette, Alessandro
Sette, Alessandro
中科院分区:
生物学2区
文献类型:
--
作者:
Peters, Bjoern;Bui, Huynh-Hoa;Frankild, Sune;Nielsen, Morten;Lundegaard, Claus;Kostem, Emrah;Basch, Derek;Lamberth, Kasper;Harndahl, Mikkel;Fleri, Ward;Wilson, Stephen S.;Sidney, John;Lund, Ole;Buus, Soren;Sette, Alessandro

文献摘要

参考文献

被引文献

相似文献

T淋巴细胞识别与主要组织相容性复合体(MHC) I类分子结合的肽是免疫监视的重要组成部分。每个MHC等位基因都有一个特征肽结合偏好,这可以在预测算法中捕获,从而允许快速扫描整个病原体蛋白质组,以寻找可能结合MHC的肽。在这里,我们公开了一组48,828个与48种不同的小鼠,人类,猕猴和黑猩猩MHC I类等位基因相关的定量肽结合亲和力测量。我们利用这些数据建立了一组基准预测,其中一种神经网络方法和两种基于矩阵的预测方法在我们的小组中广泛使用。一般来说,神经网络优于基于矩阵的预测主要是由于它在少量数据上的泛化能力。我们还从互联网上公开的工具中检索了预测结果。虽然用于生成这些预测的数据差异妨碍了直接比较,但我们确实得出结论,基于组合肽库的工具表现非常好。该数据集的透明预测评估为工具开发人员提供了一个比较新开发的预测方法的基准。此外,为了生成和评估我们自己的预测方法,我们建立了一个易于扩展的基于web的预测框架,允许专家实现的预测方法的自动并排比较。这是对工具开发人员必须自己生成参考预测的当前实践的一个进步,这可能导致低估他们不熟悉的预测方法的性能。这项工作的总体目标是提供一个透明的预测评估,使生物信息学家能够识别预测方法的有前途的特征,并为免疫学家提供有关预测工具可靠性的指导。在高等生物中,主要的组织相容性复合体(MHC) I类分子几乎存在于所有细胞表面,在那里它们向免疫系统的T淋巴细胞提供肽。这些多肽来源于细胞内表达的蛋白质,从而允许免疫系统“窥视”细胞内部以检测感染或癌细胞。存在不同的MHC分子,每个分子都具有不同的肽结合特异性。已经开发了许多算法,可以预测哪些肽与给定的MHC分子结合。这些算法被免疫学家用来扫描特定病毒的蛋白质组,寻找可能出现在感染细胞上的肽。在本文中,作者提供了定量mhc肽结合数据的大规模实验数据集。利用这个数据集,他们比较了不同的方法识别结合肽的能力。这种比较确定了人工神经网络是目前可用的最成功的肽结合预测方法。这种比较可以作为未来工具开发的基准,使生物信息学家能够记录工具开发的进展,并指导免疫学家选择良好的预测算法。
Recognition of peptides bound to major histocompatibility complex (MHC) class I molecules by T lymphocytes is an essential part of immune surveillance. Each MHC allele has a characteristic peptide binding preference, which can be captured in prediction algorithms, allowing for the rapid scan of entire pathogen proteomes for peptide likely to bind MHC. Here we make public a large set of 48,828 quantitative peptide-binding affinity measurements relating to 48 different mouse, human, macaque, and chimpanzee MHC class I alleles. We use this data to establish a set of benchmark predictions with one neural network method and two matrix-based prediction methods extensively utilized in our groups. In general, the neural network outperforms the matrix-based predictions mainly due to its ability to generalize even on a small amount of data. We also retrieved predictions from tools publicly available on the internet. While differences in the data used to generate these predictions hamper direct comparisons, we do conclude that tools based on combinatorial peptide libraries perform remarkably well. The transparent prediction evaluation on this dataset provides tool developers with a benchmark for comparison of newly developed prediction methods. In addition, to generate and evaluate our own prediction methods, we have established an easily extensible web-based prediction framework that allows automated side-by-side comparisons of prediction methods implemented by experts. This is an advance over the current practice of tool developers having to generate reference predictions themselves, which can lead to underestimating the performance of prediction methods they are not as familiar with as their own. The overall goal of this effort is to provide a transparent prediction evaluation allowing bioinformaticians to identify promising features of prediction methods and providing guidance to immunologists regarding the reliability of prediction tools. In higher organisms, major histocompatibility complex (MHC) class I molecules are present on nearly all cell surfaces, where they present peptides to T lymphocytes of the immune system. The peptides are derived from proteins expressed inside the cell, and thereby allow the immune system to “peek inside” cells to detect infections or cancerous cells. Different MHC molecules exist, each with a distinct peptide binding specificity. Many algorithms have been developed that can predict which peptides bind to a given MHC molecule. These algorithms are used by immunologists to, for example, scan the proteome of a given virus for peptides likely to be presented on infected cells. In this paper, the authors provide a large-scale experimental dataset of quantitative MHC–peptide binding data. Using this dataset, they compare how well different approaches are able to identify binding peptides. This comparison identifies an artificial neural network as the most successful approach to peptide binding prediction currently available. This comparison serves as a benchmark for future tool development, allowing bioinformaticians to document advances in tool development as well as guiding immunologists to choose good prediction algorithm.
DOI: 10.1002/eji.200425811
发表时间: 2005-08-01
影响因子: 5.4
作者:
Larsen, MV;Lundegaard, C;Nielsen, M
通讯作者: Nielsen, M
DOI: 10.1093/bioinformatics/btg424
发表时间: 2004-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bhasin, M;Raghava, GPS
通讯作者: Raghava, GPS
DOI: 10.1006/jmbi.1997.0937
发表时间: 1997-04-18
影响因子: 5.6
作者:
Gulukota, K;Sidney, J;DeLisi, C
通讯作者: DeLisi, C
DOI: 10.1007/s002510100300
发表时间: 2001-03-01
期刊: IMMUNOGENETICS
影响因子: 3.2
作者:
Nussbaum, AK;Kuttler, C;Schild, H
通讯作者: Schild, H
DOI: 10.1093/bioinformatics/bth100
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Nielsen, M;Lundegaard, C;Lund, O
通讯作者: Lund, O