Fallacy of the Unique Genome: Sequence Diversity within Single Helicobacter pylori Strains.

Fallacy of the Unique Genome: Sequence Diversity within Single Helicobacter pylori Strains.
复制标题

DOI:
10.1128/mbio.02321-16
复制
发表时间:
2017-02-21
期刊:
影响因子:
6.4
通讯作者:
Ottemann KM
Ottemann KM
中科院分区:
生物学1区
文献类型:
--
作者:
Draper JL;Hansen LM;Bernick DL;Abedrabbo S;Underwood JG;Kong N;Huang BC;Weis AM;Weimer BC;van Vliet AH;Pourmand N;Solnick JV;Karplus K;Ottemann KM

文献摘要

被引文献

相似文献

许多细菌基因组是高度可变的,但通常作为单个组装的基因组发表。跟踪细菌基因组进化的实验没有研究在给定时间点存在的变异。在这里,我们分析了小鼠传代幽门螺杆菌菌株SS 1及其亲本PMSS 1,以评估基因组内和基因组间的变异性。使用高序列覆盖深度和实验验证,我们在这些幽门螺杆菌分离株中检测到广泛的基因组可塑性,包括转座因子IS 607的移动、大小倒位、多个单核苷酸多态性和cagA拷贝数的变异。cagA基因被发现为1至4个串联拷贝位于关闭的cag岛在SS 1和PMSS 1,这种拷贝数的变化与蛋白质表达。为了深入了解小鼠适应过程中发生的变化,我们还比较了SS 1和PMSS 1,并观察到46个与基因组内变异不同的差异。最重要的是插入cagY,其编码IV型分泌系统功能所需的蛋白质。我们检测了已知影响小鼠定植的两种蛋白质编码基因的修饰,即HpaA神经氨酸乳糖结合蛋白和FutB α-1,3脂多糖(LPS)岩藻糖基转移酶,以及预测调节不同性质的基因。总之,我们的工作表明,来自单菌落的共有基因组组装的数据可能会因未能代表存在的变异性而产生误导。此外,我们表明,可以分析群体的高深度基因组测序数据,以了解细菌菌株内的正常变异。虽然众所周知,许多细菌基因组是高度可变的,但是参考、分析和公布细菌菌株的“基因组”仍然是传统的。变异性通常被降低(“仅来自单个菌落的序列”)、被忽略(“仅发布共识”)或被置于“太难”的篮子中(“原始读取数据的分析更稳健”)。现在全基因组序列经常用于评估毒力和跟踪疫情,需要更好地了解单个菌株中存在的基线基因组变异。在这里,我们描述的变异性,在典型的工作股票和菌落的病原体幽门螺杆菌模型菌株SS 1和PMSS 1所揭示的使用高覆盖率的配偶对下一代测序(NGS),并确认了传统的实验室技术。这项工作表明,依赖于作为细菌菌株的“基因组”的共有组装可能会产生误导。
Many bacterial genomes are highly variable but nonetheless are typically published as a single assembled genome. Experiments tracking bacterial genome evolution have not looked at the variation present at a given point in time. Here, we analyzed the mouse-passaged Helicobacter pylori strain SS1 and its parent PMSS1 to assess intra- and intergenomic variability. Using high sequence coverage depth and experimental validation, we detected extensive genome plasticity within these H. pylori isolates, including movement of the transposable element IS607, large and small inversions, multiple single nucleotide polymorphisms, and variation in cagA copy number. The cagA gene was found as 1 to 4 tandem copies located off the cag island in both SS1 and PMSS1; this copy number variation correlated with protein expression. To gain insight into the changes that occurred during mouse adaptation, we also compared SS1 and PMSS1 and observed 46 differences that were distinct from the within-genome variation. The most substantial was an insertion in cagY, which encodes a protein required for a type IV secretion system function. We detected modifications in genes coding for two proteins known to affect mouse colonization, the HpaA neuraminyllactose-binding protein and the FutB α-1,3 lipopolysaccharide (LPS) fucosyltransferase, as well as genes predicted to modulate diverse properties. In sum, our work suggests that data from consensus genome assemblies from single colonies may be misleading by failing to represent the variability present. Furthermore, we show that high-depth genomic sequencing data of a population can be analyzed to gain insight into the normal variation within bacterial strains. Although it is well known that many bacterial genomes are highly variable, it is nonetheless traditional to refer to, analyze, and publish “the genome” of a bacterial strain. Variability is usually reduced (“only sequence from a single colony”), ignored (“just publish the consensus”), or placed in the “too-hard” basket (“analysis of raw read data is more robust”). Now that whole-genome sequences are regularly used to assess virulence and track outbreaks, a better understanding of the baseline genomic variation present within single strains is needed. Here, we describe the variability seen in typical working stocks and colonies of pathogen Helicobacter pylori model strains SS1 and PMSS1 as revealed by use of high-coverage mate pair next-generation sequencing (NGS) and confirmed by traditional laboratory techniques. This work demonstrates that reliance on a consensus assembly as “the genome” of a bacterial strain may be misleading.