Data-driven insights into deletions of Mycobacterium tuberculosis complex chromosomal DR region using spoligoforests.

Data-driven insights into deletions of Mycobacterium tuberculosis complex chromosomal DR region using spoligoforests.
复制标题

使用 spoligoforests 对结核分枝杆菌复合体染色体 DR 区域的删除进行数据驱动的见解。

DOI:
10.1109/bibm.2011.64
复制
发表时间:
2011
期刊:
Proceedings. IEEE International Conference on Bioinformatics and Biomedicine
影响因子:
--
通讯作者:
Bennett,KristinP
Bennett,KristinP
中科院分区:
--
文献类型:
--
作者:
Ozcaglar,Cagri;Shabbeer,Amina;Kurepina,Natalia;Yener,Bülent;Bennett,KristinP

文献摘要

相似文献

结核分枝杆菌复合群(MTBC)的生物标志物随时间发生突变。在MTBC的生物标志物中,间隔寡核苷酸型(spoligotype)和分枝杆菌散布重复单位(MIRU)模式通常用于对临床MTBC菌株进行基因分型。在这项研究中,我们提出了一个进化模型的spoligotype重排使用MIRU模式,以消除歧义的spoligotypes的祖先,在一个大型的患者数据集从美国疾病控制和预防中心(CDC)。基于连续缺失假设和罕见的观察收敛进化,我们首先产生最简约的森林spoligotypes,称为spoligoforest,使用三个遗传距离的措施。spoligoforest的拓扑属性和在每个菌株的直接重复(DR)位点的变化的数量的分析揭示了有趣的DR区域中的缺失的属性。首先,我们将我们的突变模型与现有的spoligotypes突变模型进行比较,发现我们的突变模型产生的谱系内突变事件与其他模型一样多,分离精度略高。其次,基于我们的突变模型,后代spoligotypes的数量遵循幂律分布。第三,与以前的研究相反,幂律分布不符合突变长度频率。最后,在连续的DR基因座的突变事件的总数遵循双峰分布,这导致在DR区域的较短的缺失的积累。这两种模式是间隔区13和40,它们是染色体重排的热点。双峰分布中的变化点是间隔区34,其在大多数MTBC菌株中不存在。这种双峰分离导致较短缺失的积累,这解释了为什么幂律分布与突变长度频率不合理。
Biomarkers of Mycobacterium tuberculosis complex (MTBC) mutate over time. Among the biomarkers of MTBC, spacer oligonucleotide type (spoligotype) and Mycobacterium Interspersed Repetitive Unit (MIRU) patterns are commonly used to genotype clinical MTBC strains. In this study, we present an evolution model of spoligotype rearrangements using MIRU patterns to disambiguate the ancestors of spoligotypes, in a large patient dataset from the United States Centers for Disease Control and Prevention (CDC). Based on the contiguous deletion assumption and rare observation of convergent evolution, we first generate the most parsimonious forest of spoligotypes, called a spoligoforest, using three genetic distance measures. An analysis of topological attributes of the spoligoforest and number of variations at the direct repeat (DR) locus of each strain reveals interesting properties of deletions in the DR region. First, we compare our mutation model to existing mutation models of spoligotypes and find that our mutation model produces as many within-lineage mutation events as other models, with slightly higher segregation accuracy. Second, based on our mutation model, the number of descendant spoligotypes follows a power law distribution. Third, contrary to prior studies, the power law distribution does not plausibly fit to the mutation length frequency. Finally, the total number of mutation events at consecutive DR loci follows a bimodal distribution, which results in accumulation of shorter deletions in the DR region. The two modes are spacers 13 and 40, which are hotspots for chromosomal rearrangements. The change point in the bimodal distribution is spacer 34, which is absent in most MTBC strains. This bimodal separation results in accumulation of shorter deletions, which explains why a power law distribution is not a plausible fit to the mutation length frequency.