Maximum entropy models for antibody diversity

Maximum entropy models for antibody diversity
复制标题

DOI:
10.1073/pnas.1001705107
复制
发表时间:
2010-03-23
影响因子:
11.1
通讯作者:
Callan, Curtis G., Jr.
Callan, Curtis G., Jr.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Mora, Thierry;Walczak, Aleksandra M.;Callan, Curtis G., Jr.

文献摘要

被引文献

相似文献

病原体的识别依赖于表现出巨大多样性的蛋白质家族。在这里,我们构建了序列库的最大熵模型,以最近的实验为基础,这些实验提供了斑马鱼 IgM 序列的近乎详尽的采样。这些模型仅基于残基位置之间的成对相关性,但正确捕获了指令集的高阶统计特性。通过利用这些模型对统计物理问题的解释,我们对序列集合的集体属性做出了一些预测:序列的分布服从齐普夫定律,全部分解成几个簇,并且由于相关性而存在巨大的多样性限制。这些预测与在每个位点独立进行氨基酸取代的模型完全不一致,并且与数据非常一致。我们的结果表明,抗体多样性不受基因组中编码的序列的限制,并且可能反映了对抗原挑战的快速适应。这种方法应该适用于其他蛋白质家族的整体特性的研究。
Recognition of pathogens relies on families of proteins showing great diversity. Here we construct maximum entropy models of the sequence repertoire, building on recent experiments that provide a nearly exhaustive sampling of the IgM sequences in zebrafish. These models are based solely on pairwise correlations between residue positions but correctly capture the higher order statistical properties of the repertoire. By exploiting the interpretation of these models as statistical physics problems, we make several predictions for the collective properties of the sequence ensemble: The distribution of sequences obeys Zipf's law, the repertoire decomposes into several clusters, and there is a massive restriction of diversity because of the correlations. These predictions are completely inconsistent with models in which amino acid substitutions are made independently at each site and are in good agreement with the data. Our results suggest that antibody diversity is not limited by the sequences encoded in the genome and may reflect rapid adaptation to antigenic challenges. This approach should be applicable to the study of the global properties of other protein families.