Optimizing hybrid assembly of next-generation sequence data from Enterococcus faecium: a microbe with highly divergent genome.

Optimizing hybrid assembly of next-generation sequence data from Enterococcus faecium: a microbe with highly divergent genome.
复制标题

优化屎肠球菌下一代序列数据的杂交组装:一种具有高度差异基因组的微生物

DOI:
10.1186/1752-0509-6-s3-s21
复制
发表时间:
2012
影响因子:
--
通讯作者:
Li X
Li X
中科院分区:
生物2区
文献类型:
--
作者:
Wang Y;Yu Y;Pan B;Hao P;Li Y;Shao Z;Xu X;Li X

文献摘要

被引文献

相似文献

背景 细菌基因组测序成为研究病原体毒力和密切相关菌株之间系统发育关系的重要方法。屎肠球菌成为一种重要的医院病原体,通常与医院对常见抗生素的耐药性有关。由于基因内容差异较大,对高通量、短读长的下一代测序(NGS)技术提出了挑战。本研究旨在调查 NGS 技术的特性和系统偏差,并使用 NGS 数据组合评估影响混合组装结果的关键参数。 结果 使用三种不同的 NGS 平台对医院的屎肠球菌菌株进行测序:454 GS-FLX、Illumina GAIIx 和 ABI SOLiD4.0,覆盖深度约为 28、500 和 400 倍。我们构建了一个管道,将每个 NGS 数据中的重叠群合并到混合组件中。结果表明,每个 NGS 组件的连续性都存在上限,无法通过简单地增加数据覆盖深度来克服。每种 NGS 技术都表现出一些内在特性,即碱基识别错误、系统偏差等。每个 NGS 组装的间隙和低覆盖区域与较低的 GC 含量相关。为了优化混合组装方法,我们使用不同数量和不同组合的 NGS 数据进行测试,并获得了组装连续性的最佳条件。我们还首次表明,SOLiD 数据与其他类型的 NGS 数据结合使用混合方法可以帮助大大改进屎肠球菌基因组的组装。 结论 当前的研究解决了如何利用当今最先进的测序技术最有效地构建完整的微生物基因组的难题。我们对每种 NGS 技术的序列数据和基因组组装进行了表征,测试了结合 NGS 数据的混合组装条件,并获得了实现最具成本效益组装的优化参数。我们的研究帮助形成了一些指导其他微生物基因组工作的指南,因此具有重要的实际意义。
Background Sequencing of bacterial genomes became an essential approach to study pathogen virulence and the phylogenetic relationship among close related strains. Bacterium Enterococcus faecium emerged as an important nosocomial pathogen that were often associated with resistance to common antibiotics in hospitals. With highly divergent gene contents, it presented a challenge to the next generation sequencing (NGS) technologies featuring high-throughput and shorter read-length. This study was designed to investigate the properties and systematic biases of NGS technologies and evaluate critical parameters influencing the outcomes of hybrid assemblies using combinations of NGS data. Results A hospital strain of E. faecium was sequenced using three different NGS platforms: 454 GS-FLX, Illumina GAIIx, and ABI SOLiD4.0, to approximately 28-, 500-, and 400-fold coverage depth. We built a pipeline that merged contigs from each NGS data into hybrid assemblies. The results revealed that each single NGS assembly had a ceiling in continuity that could not be overcome by simply increasing data coverage depth. Each NGS technology displayed some intrinsic properties, i.e. base calling error, systematic bias, etc. The gaps and low coverage regions of each NGS assembly were associated with lower GC contents. In order to optimize the hybrid assembly approach, we tested with varying amount and different combination of NGS data, and obtained optimal conditions for assembly continuity. We also, for the first time, showed that SOLiD data could help make much improved assemblies of E. faecium genome using the hybrid approach when combined with other type of NGS data. Conclusions The current study addressed the difficult issue of how to most effectively construct a complete microbial genome using today's state of the art sequencing technologies. We characterized the sequence data and genome assembly from each NGS technologies, tested conditions for hybrid assembly with combinations of NGS data, and obtained optimized parameters for achieving most cost-efficiency assembly. Our study helped form some guidelines to direct genomic work on other microorganisms, thus have important practical implications.