Whole genome sequence and manual annotation of Clostridium autoethanogenum, an industrially relevant bacterium.

Whole genome sequence and manual annotation of Clostridium autoethanogenum, an industrially relevant bacterium.
复制标题

DOI:
10.1186/s12864-015-2287-5
复制
发表时间:
2015-12-21
期刊:
影响因子:
4.4
通讯作者:
Minton NP
Minton NP
中科院分区:
生物学2区
文献类型:
--
作者:
Humphreys CM;McLean S;Schatschneider S;Millat T;Henstra AM;Annan FJ;Breitkopf R;Pander B;Piatek P;Rowe P;Wichlacz AT;Woods C;Norman R;Blom J;Goesman A;Hodgman C;Barrett D;Thomas NR;Winzer K;Minton NP

文献摘要

被引文献

相似文献

自产乙酸梭菌是一种产乙酸菌,能够从合成气中存在的C1气体中生产高价值的商品化学品和生物燃料。这种常见的工业废气可以作为细菌的唯一能源和碳源,通过还原的乙酰-辅酶A(Wood-Ljugdahl)途径将低价值气体成分转化为细胞构建块和工业相关产品。目前的研究工作集中在通过合成生物学方法加强和扩大这种有机体中的产物形成。然而,对代谢建模和定向途径工程至关重要的是可靠的和全面注释的基因组序列。我们使用Illumina MiSeq技术对DSM10061株自乙酸梭菌进行了下一代测序,发现与已公布的完成序列(NCBI:Gca_000484505.1)相比有243个单核苷酸差异,其中59.1%位于编码区。Sanger测序证实了这些变异,随后的分析表明,这些差异是已发表基因组中的测序错误,而不是真正的单核苷酸多态。观察到90%以上发生在长度大于4个核苷酸的均聚物区内,证实了这一点。还有人注意到,许多包含这些测序错误的基因在已发表的封闭基因组中被注释为编码含有移码突变的蛋白质(18例),或者尽管编码框架包含终止密码子,但仍被注释,如果终止密码子是真的,将严重阻碍有机体的生存能力。此外,我们还完成了全面的人工整理,以减少因在相关物种中连续使用自动标注管道而出现的标注错误。结果,不同的功能被赋予基因产品或以前的功能注释,因为在不同的场合缺少证据而被拒绝。我们提出了一个修订的人工挑选的自产乙酸梭菌DSM10061的全基因组序列,它为严重依赖注释准确性的基因组规模模型提供了可靠的信息,并代表了朝着操纵和代谢模型这一工业相关的乙酸原迈出的重要一步。本文的在线版本(doi:10.1186/s12864-0152287-5)包含补充材料,授权用户可以使用。
Clostridium autoethanogenum is an acetogenic bacterium capable of producing high value commodity chemicals and biofuels from the C1 gases present in synthesis gas. This common industrial waste gas can act as the sole energy and carbon source for the bacterium that converts the low value gaseous components into cellular building blocks and industrially relevant products via the action of the reductive acetyl-CoA (Wood-Ljungdahl) pathway. Current research efforts are focused on the enhancement and extension of product formation in this organism via synthetic biology approaches. However, crucial to metabolic modelling and directed pathway engineering is a reliable and comprehensively annotated genome sequence. We performed next generation sequencing using Illumina MiSeq technology on the DSM10061 strain of Clostridium autoethanogenum and observed 243 single nucleotide discrepancies when compared to the published finished sequence (NCBI: GCA_000484505.1), with 59.1 % present in coding regions. These variations were confirmed by Sanger sequencing and subsequent analysis suggested that the discrepancies were sequencing errors in the published genome not true single nucleotide polymorphisms. This was corroborated by the observation that over 90 % occurred within homopolymer regions of greater than 4 nucleotides in length. It was also observed that many genes containing these sequencing errors were annotated in the published closed genome as encoding proteins containing frameshift mutations (18 instances) or were annotated despite the coding frame containing stop codons, which if genuine, would severely hinder the organism’s ability to survive. Furthermore, we have completed a comprehensive manual curation to reduce errors in the annotation that occur through serial use of automated annotation pipelines in related species. As a result, different functions were assigned to gene products or previous functional annotations rejected because of missing evidence in various occasions. We present a revised manually curated full genome sequence for Clostridium autoethanogenum DSM10061, which provides reliable information for genome-scale models that rely heavily on the accuracy of annotation, and represents an important step towards the manipulation and metabolic modelling of this industrially relevant acetogen. The online version of this article (doi:10.1186/s12864-015-2287-5) contains supplementary material, which is available to authorized users.