repgenHMM: a dynamic programming tool to infer the rules of immune receptor generation from sequence data.

repgenHMM: a dynamic programming tool to infer the rules of immune receptor generation from sequence data.
复制标题

DOI:
10.1093/bioinformatics/btw112
复制
发表时间:
2016-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Walczak AM
Walczak AM
中科院分区:
其他
文献类型:
--
作者:
Elhanati Y;Marcou Q;Mora T;Walczak AM

文献摘要

被引文献

相似文献

动机:免疫库的多样性最初是由早期T和B细胞发育过程中受体基因的随机重排产生的。重排的情况是由随机事件组成的基因模板的选择,碱基对删除和插入概率分布描述。并非所有的情况都是同样可能的,相同的受体序列可以通过几种不同的方式获得。量化这些重排的分布是研究免疫系统多样性的基本基线。从受体序列推断分布的性质是一个计算困难的问题,需要枚举每个采样受体序列的每一种可能的情况。结果:我们提出了一个隐马尔可夫模型,它占所有可能的情况下,可以产生的受体序列。我们开发并实现了一种基于Baum-Welch算法的方法,该方法可以有效地推断重排过程中不同事件的参数。我们在T细胞受体的α链和β链的序列数据上测试了我们的软件工具。为了测试我们的算法的有效性,我们还生成了由已知模型产生的合成序列,并证实其参数可以准确地从序列中推断出来。推断的模型可用于生成合成序列,以计算任何受体序列的生成概率,以及库的理论多样性。我们估计这种多样性是人类T细胞。该模型为研究免疫库的选择和动态提供了一个基线。可用性和实现:源代码和样本序列文件可在https://bitbucket.org/yuvalel/repgenhmm/downloads上获得。联系方式:elhanati@lpt.ens.fr或tmora@lps.ens.fr或awalczak@lpt.ens.fr
Motivation: The diversity of the immune repertoire is initially generated by random rearrangements of the receptor gene during early T and B cell development. Rearrangement scenarios are composed of random events—choices of gene templates, base pair deletions and insertions—described by probability distributions. Not all scenarios are equally likely, and the same receptor sequence may be obtained in several different ways. Quantifying the distribution of these rearrangements is an essential baseline for studying the immune system diversity. Inferring the properties of the distributions from receptor sequences is a computationally hard problem, requiring enumerating every possible scenario for every sampled receptor sequence. Results: We present a Hidden Markov model, which accounts for all plausible scenarios that can generate the receptor sequences. We developed and implemented a method based on the Baum–Welch algorithm that can efficiently infer the parameters for the different events of the rearrangement process. We tested our software tool on sequence data for both the alpha and beta chains of the T cell receptor. To test the validity of our algorithm, we also generated synthetic sequences produced by a known model, and confirmed that its parameters could be accurately inferred back from the sequences. The inferred model can be used to generate synthetic sequences, to calculate the probability of generation of any receptor sequence, as well as the theoretical diversity of the repertoire. We estimate this diversity to be for human T cells. The model gives a baseline to investigate the selection and dynamics of immune repertoires. Availability and implementation: Source code and sample sequence files are available at https://bitbucket.org/yuvalel/repgenhmm/downloads. Contact: elhanati@lpt.ens.fr or tmora@lps.ens.fr or awalczak@lpt.ens.fr