A critical comparison of technologies for a plant genome sequencing project

A critical comparison of technologies for a plant genome sequencing project
复制标题

DOI:
10.1101/201830
复制
发表时间:
2017-10
期刊:
影响因子:
9.2
通讯作者:
Pirita Paajanen;George Kettleborough;E. López-Girona;M. Giolai;D. Heavens;David Baker;Ashleigh Lister-Ashleigh-Lis
Pirita Paajanen;George Kettleborough;E. López-Girona;M. Giolai;D. Heavens;David Baker;Ashleigh Lister-Ashleigh-Lis
中科院分区:
生物学2区
文献类型:
--
作者:
Pirita Paajanen;George Kettleborough;E. López-Girona;M. Giolai;D. Heavens;David Baker;Ashleigh Lister-Ashleigh-Lis

文献摘要

被引文献

相似文献

高质量的模式生物基因组序列是许多研究的重要起点。旧的基于克隆的方法既慢又贵,而更快、更便宜的短只读程序集可能是不完整的和高度碎片化的,这就降低了它们的有用性。在过去的几年里,许多新的基因组组装技术相继问世。这些新技术和新算法通常以微生物基因组为基准,如果规模适当,则以人类基因组为基准。然而,植物基因组可能比人类更重复、更大,植物生物学使获得高质量、不受污染的DNA变得困难。反映了它们具有挑战性的性质,我们观察到植物基因组组装统计数据通常比脊椎动物差。在这里,我们比较了Illumina Short Read、PacBio Long Read、10倍基因组连锁Reads、Dovetail Hi-C和BioNano Genome光学图谱,单独和组合在一起,在生产马铃薯种疣状链霉菌高质量的远程基因组组合方面。我们对组装的完整性和准确性以及DNA、计算要求和测序成本进行基准测试。我们希望我们的结果将有助于其他基因组计划,并希望这些数据集将用于汇编算法开发人员的基准测试。
A high quality genome sequence of your model organism is an essential starting point for many studies. Old clone based methods are slow and expensive, whereas faster, cheaper short read only assemblies can be incomplete and highly fragmented, which minimises their usefulness. The last few years have seen the introduction of many new technologies for genome assembly. These new technologies and new algorithms are typically benchmarked on microbial genomes or, if they scale appropriately, human. However, plant genomes can be much more repetitive and larger than human, and plant biology makes obtaining high quality DNA free from contaminants difficult. Reflecting their challenging nature we observe that plant genome assembly statistics are typically poorer than for vertebrates. Here we compare Illumina short read, PacBio long read, 10x Genomics linked reads, Dovetail Hi-C and BioNano Genomics optical maps, singly and combined, in producing high quality long range genome assemblies of the potato species S. verrucosum. We benchmark the assemblies for completeness and accuracy, as well as DNA, compute requirements and sequencing costs. We expect our results will be helpful to other genome projects, and that these datasets will be used in benchmarking by assembly algorithm developers.