Consistency of VDJ Rearrangement and Substitution Parameters Enables Accurate B Cell Receptor Sequence Annotation

Consistency of VDJ Rearrangement and Substitution Parameters Enables Accurate B Cell Receptor Sequence Annotation
复制标题

DOI:
10.1371/journal.pcbi.1004409
复制
发表时间:
2016-01-01
影响因子:
4.3
通讯作者:
Matsen, Frederick A.
Matsen, Frederick A.
中科院分区:
生物学2区
文献类型:
--
作者:
Ralph, Duncan K.;Matsen, Frederick A.

文献摘要

被引文献

相似文献

VDJ重排和体细胞超突变共同作用产生抗体编码B细胞受体(BCR)序列,用于显著多样性的抗原。现在可以高通量地对这些BCR进行测序;对这些序列的分析为抗体的发展带来了新的见解,特别是针对HIV和流感的广泛中和抗体。此类序列分析的基本步骤是将每个碱基注释为来自V、D或J基因中的特定一个,或来自N添加(又名非模板化插入)。以前的工作已经使用简单的参数分布模型从状态到状态的隐马尔可夫模型(HMM)的VDJ重组的过渡,并假设突变发生通过相同的过程跨站点。然而,已经观察到密码子框架和其他效应违反了这些编码序列的参数假设,这表明对重组过程建模的非参数方法可能是有用的。在我们的论文中,我们发现,实际上大型的现代数据集建议使用参数丰富的每个等位基因的分类分布的HMM转移概率和每个等位基因每个位置的突变概率的模型,并使用这样的模型进行推理导致显着改善的结果。我们提出了一个准确和高效的BCR序列注释软件包,使用一种新的HMM“因子分解”策略。这个软件包称为partis(https://github.com/psathyrella/partis/),构建在一个新的通用HMM编译器上,该编译器可以在给定HMM的简单文本描述的情况下执行有效的推理。
VDJ rearrangement and somatic hypermutation work together to produce antibody-coding B cell receptor (BCR) sequences for a remarkable diversity of antigens. It is now possible to sequence these BCRs in high throughput; analysis of these sequences is bringing new insight into how antibodies develop, in particular for broadly-neutralizing antibodies against HIV and influenza. A fundamental step in such sequence analysis is to annotate each base as coming from a specific one of the V, D, or J genes, or from an N-addition (a.k.a. non-templated insertion). Previous work has used simple parametric distributions to model transitions from state to state in a hidden Markov model (HMM) of VDJ recombination, and assumed that mutations occur via the same process across sites. However, codon frame and other effects have been observed to violate these parametric assumptions for such coding sequences, suggesting that a non-parametric approach to modeling the recombination process could be useful. In our paper, we find that indeed large modern data sets suggest a model using parameter-rich per-allele categorical distributions for HMM transition probabilities and per-allele-per-position mutation probabilities, and that using such a model for inference leads to significantly improved results. We present an accurate and efficient BCR sequence annotation software package using a novel HMM "factorization" strategy. This package, called partis (https://github.com/psathyrella/partis/), is built on a new general-purpose HMM compiler that can perform efficient inference given a simple text description of an HMM.