Identifying highly mutated IGHD genes in the junctions of rearranged human immunoglobulin heavy chain genes

Identifying highly mutated IGHD genes in the junctions of rearranged human immunoglobulin heavy chain genes
复制标题

DOI:
10.1016/j.jim.2007.04.011
复制
发表时间:
2007-07-31
影响因子:
2.2
通讯作者:
Collins, Andrew M.
Collins, Andrew M.
中科院分区:
医学4区
文献类型:
--
作者:
Jackson, Katherine J. L.;Gaeta, Bruno A.;Collins, Andrew M.

文献摘要

被引文献

相似文献

人免疫球蛋白重链内IGHD基因的可靠鉴定具有挑战性,高达三分之一的重排没有可鉴定的IGHD基因。短的、突变的IGHD基因通常被认为与它们周围的非模板编码核苷酸的N区是不可区分的。在这项研究中,我们已经表征了N-区域,证明了核苷酸组成的偏见在添加过程中的重要性,包括形成均聚物束。然后,我们使用模拟的方法来确定高度突变的IGHD基因之间的连接核苷酸的错误识别的可能性。这些可能性为鉴定突变的D-REGION提供了一般规则,并表明可以以低错误风险鉴定具有多达10个突变的较长D-REGION(>25个核苷酸)。具有多达四个突变的较短D-REGION(> 16个核苷酸)也是可鉴定的。不同对齐的可靠性取决于结的长度(N-REGION和D-REGION的组合)。提供的数据可以指导连接长度为5至50个核苷酸的序列的比对,包括两个D区可能性之间的明确选择。使用这种基于免疫学的方法来比对IGHD基因将提高免疫球蛋白序列划分的可靠性,这反过来又将促进对免疫球蛋白库多样性的许多过程的研究。(C)2007 Elsevier B. V.保留所有权利。
The reliable identification of IGHD genes within human immunoglobulin heavy chains is challenging with up to one third of rearrangements having no identifiable IGHD gene. The short, mutated IGHD genes are generally assumed to be indistinguishable from the N-REGIONS of non-template encoded nucleotides that surround them. In this study we have characterised N-REGIONS, demonstrating the importance of nucleotide composition biases in the addition process, including the formation of homopolymer tracts. We then use a simulation approach to determine the likelihood of misidentification of highly mutated IGHD genes among the JUNCTION nucleotides. These likelihoods provide general rules for the identification of mutated D-REGIONs, and suggest that longer D-REGIONs (>25 nucleotides) with as many as ten mutations can be identified with a low risk of error. Shorter D-REGIONs (> 16 nucleotides) with as many as four mutations are also identifiable. The reliability of different alignments is dependent upon the junction length (combined N-REGIONs and D-REGION). Data is presented that can guide the alignment of sequences with junction lengths from 5 to 50 nucleotides, including explicit selection between two D-REGION possibilities. The use of such a statistically-based approach to the alignment of IGHD genes will improve the reliability of the partitioning of immunoglobulin sequences, and this in turn will facilitate the study of the many processes that contribute to the diversity of the immunoglobulin repertoire. (C) 2007 Elsevier B.V. All rights reserved.