CSA: A high-throughput chromosome-scale assembly pipeline for vertebrate genomes

CSA: A high-throughput chromosome-scale assembly pipeline for vertebrate genomes
复制标题

DOI:
10.1093/gigascience/giaa034
复制
发表时间:
2020-05
期刊:
影响因子:
9.2
通讯作者:
Heiner Kuhl;Ling Li;S. Wuertz;M. Stöck;Xu-Fang Liang;C. Klopp
Heiner Kuhl;Ling Li;S. Wuertz;M. Stöck;Xu-Fang Liang;C. Klopp
中科院分区:
生物学2区
文献类型:
--
作者:
Heiner Kuhl;Ling Li;S. Wuertz;M. Stöck;Xu-Fang Liang;C. Klopp

文献摘要

相似文献

摘要背景用于长阅读组装的易于使用和快速的生物信息学管道超出了重叠群水平,以从原始数据生成高度连续的染色体规模的基因组的情况仍然很少。结果染色体规模汇编器(CSA)是一种新型的计算高效的生物信息学流水线,填补了这一空白。CSA将来自支架组装(例如,Hi-C或10X基因组)或甚至来自分散的参考基因组的信息整合到组装过程中。当CSA执行染色体大小的支架的自动组装时,我们将其性能与最先进的参考基因组进行基准测试,即传统地使用多个单独的组装工具和手动管理以繁琐的方式构建。CSA通过支架、局部重组和缝合来增加重叠群长度。在某些数据集上,初始重叠群N50可能增加到4.5倍。对于较小的脊椎动物基因组,使用低成本的高端台式计算机可以在12小时内完成染色体规模的组装。哺乳动物基因组可以在16小时内在计算机服务器上处理。使用鱼类、鸟类和哺乳动物的不同参考基因组,我们证明了CSA仅根据长时间阅读的数据和基因组比较来计算染色体规模的组装。即使是分裂基因组的重叠群水平的草稿组装也有助于重建染色体规模的序列。CSA还能够组装超长读数。结论CsA可加快和简化大型家系水平脊椎动物基因组计划的染色体水平组装,并显著降低成本。
Abstract Background Easy-to-use and fast bioinformatics pipelines for long-read assembly that go beyond the contig level to generate highly continuous chromosome-scale genomes from raw data remain scarce. Result Chromosome-Scale Assembler (CSA) is a novel computationally highly efficient bioinformatics pipeline that fills this gap. CSA integrates information from scaffolded assemblies (e.g., Hi-C or 10X Genomics) or even from diverged reference genomes into the assembly process. As CSA performs automated assembly of chromosome-sized scaffolds, we benchmark its performance against state-of-the-art reference genomes, i.e., conventionally built in a laborious fashion using multiple separate assembly tools and manual curation. CSA increases the contig lengths using scaffolding, local re-assembly, and gap closing. On certain datasets, initial contig N50 may be increased up to 4.5-fold. For smaller vertebrate genomes, chromosome-scale assemblies can be achieved within 12 h using low-cost, high-end desktop computers. Mammalian genomes can be processed within 16 h on compute-servers. Using diverged reference genomes for fish, birds, and mammals, we demonstrate that CSA calculates chromosome-scale assemblies from long-read data and genome comparisons alone. Even contig-level draft assemblies of diverged genomes are helpful for reconstructing chromosome-scale sequences. CSA is also capable of assembling ultra-long reads. Conclusions CSA can speed up and simplify chromosome-level assembly and significantly lower costs of large-scale family-level vertebrate genome projects.