Insights from the revised complete genome sequences of Acinetobacter baumannii strains AB307-0294 and ACICU belonging to global clones 1 and 2

Insights from the revised complete genome sequences of Acinetobacter baumannii strains AB307-0294 and ACICU belonging to global clones 1 and 2
复制标题

DOI:
10.1099/mgen.0.000298
复制
发表时间:
2019-10-01
期刊:
影响因子:
3.9
通讯作者:
Hall, Ruth M.
Hall, Ruth M.
中科院分区:
生物学2区
文献类型:
--
作者:
Hamidian, Mohammad;Wick, Ryan R.;Hall, Ruth M.

文献摘要

被引文献

相似文献

鲍曼不动杆菌全球克隆1 (AB307-0294)于1994年在美国被发现,全球克隆2 (GC2)分离物ACICU于2005年在意大利被发现,它们是首批被完全测序的鲍曼不动杆菌分离物。AB307-0294对大多数抗生素敏感,并已在许多遗传学研究中使用,ACICU属于罕见的GC2谱系。完整的基因组序列,最初使用454 pyrosequencing技术确定,已知会产生测序错误,使用Illumina MiSeq和MinION (Oxford Nanopore Technologies)技术以及使用Unicycler生成的混合组装重新确定。将新获得的高质量基因组与先前的454个测序版本进行比较,发现了影响蛋白质编码序列(CDS)特征的大量核苷酸差异,并首次在ACICU中正确地解析了长且高度重复的bap和blp1基因的序列。比较原基因组和修订基因组的注释,发现蛋白质CDS特征存在大量差异,强调了序列错误对蛋白质序列预测和核心基因确定的影响。在修改后的基因组中,平均有400个预测的CDS更长或更短,大约200个CDS特征不再存在。
The Acinetobacter baumannii global clone 1 isolate AB307-0294, recovered in the USA in 1994, and the global clone 2 (GC2) isolate ACICU, isolated in 2005 in Italy, were among the first A. baumannii isolates to be completely sequenced. AB307-0294 is susceptible to most antibiotics and has been used in many genetic studies, and ACICU belongs to a rare GC2 lineage. The complete genome sequences, originally determined using 454 pyrosequencing technology, which is known to generate sequencing errors, were re-determined using Illumina MiSeq and MinION (Oxford Nanopore Technologies) technologies and a hybrid assembly generated using Unicycler. Comparison of the resulting new high-quality genomes to the earlier 454-sequenced versions identified a large number of nucleotide differences affecting protein coding sequence (CDS) features, and allowed the sequences of the long and highly repetitive bap and blp1 genes to be properly resolved for the first time in ACICU. Comparisons of the annotations of the original and revised genomes revealed a large number of differences in the protein CDS features, underlining the impact of sequence errors on protein sequence predictions and core gene determination. On average, 400 predicted CDSs were longer or shorter in the revised genomes and about 200 CDS features were no longer present.