A large peptidome dataset improves HLA class I epitope prediction across most of the human population

A large peptidome dataset improves HLA class I epitope prediction across most of the human population
复制标题

DOI:
10.1038/s41587-019-0322-9
复制
发表时间:
2020-02-01
影响因子:
46.9
通讯作者:
Keskin, Derin B.
Keskin, Derin B.
中科院分区:
工程技术1区
文献类型:
--
作者:
Sarkizova, Siranush;Klaeger, Susan;Keskin, Derin B.

文献摘要

被引文献

相似文献

人类白细胞抗原表位的预测对肿瘤免疫治疗和疫苗的发展具有重要意义。然而,目前的预测算法预测能力有限,部分原因是它们没有在覆盖广泛的HLA等位基因的高质量表位数据集上进行训练。为了能够在很大一部分人群中预测内源性HLAI类相关多肽,我们使用质谱仪分析了从95个HLA-A、-B、-C和-G单等位基因细胞系洗脱出来的185,000个多肽。我们确定了每个HLA等位基因的典型多肽基序,等位基因之间唯一和共享的结合亚基,以及与不同肽长度相关的不同基序。通过将这些数据与转录物丰度和多肽处理相结合,我们开发了HLAthena,为内源性多肽呈现提供了特定等位基因和长度以及泛等位基因全长预测模型。与现有工具相比,这些模型预测内源性人类白细胞抗原I类相关配体的阳性预测值提高了1.5倍,并正确识别了在11个患者来源的肿瘤细胞系中实验观察到的75%的人类白细胞抗原结合肽。
Prediction of HLA epitopes is important for the development of cancer immunotherapies and vaccines. However, current prediction algorithms have limited predictive power, in part because they were not trained on high-quality epitope datasets covering a broad range of HLA alleles. To enable prediction of endogenous HLA class I-associated peptides across a large fraction of the human population, we used mass spectrometry to profile >185,000 peptides eluted from 95 HLA-A, -B, -C and -G mono-allelic cell lines. We identified canonical peptide motifs per HLA allele, unique and shared binding submotifs across alleles and distinct motifs associated with different peptide lengths. By integrating these data with transcript abundance and peptide processing, we developed HLAthena, providing allele-and-length-specific and pan-allele-pan-length prediction models for endogenous peptide presentation. These models predicted endogenous HLA class I-associated ligands with 1.5-fold improvement in positive predictive value compared with existing tools and correctly identified >75% of HLA-bound peptides that were observed experimentally in 11 patient-derived tumor cell lines.