GAGE: A critical evaluation of genome assemblies and assembly algorithms

GAGE: A critical evaluation of genome assemblies and assembly algorithms
复制标题

DOI:
10.1101/gr.131383.111
复制
发表时间:
2012-03-01
期刊:
影响因子:
7
通讯作者:
Yorke, James A.
Yorke, James A.
中科院分区:
生物学1区
文献类型:
--
作者:
Salzberg, Steven L.;Phillippy, Adam M.;Yorke, James A.

文献摘要

被引文献

相似文献

新的测序技术极大地改变了全基因组测序的格局,使科学家能够启动许多项目来解码以前未测序的生物体的基因组。成本最低的技术可以在短短几天内对大多数物种(包括哺乳动物)进行深度覆盖。其中一个项目生成的序列数据由数百万或数十亿个短 DNA 序列(读数)组成,长度范围为 50 到 150 it。然后,在开始大多数基因组分析之前,必须重新组装这些序列。不幸的是,基因组组装仍然是一个非常困难的问题,较短的读数和不可靠的远程连接信息使这变得更加困难。在这项研究中,我们在四个不同的短读长数据集上评估了几种领先的从头组装算法,所有数据集均由 Illumina 测序仪生成。我们的结果描述了不同组装程序的相对性能以及组装难度的其他显着差异,这些差异似乎是基因组本身固有的。三个总体结论是显而易见的:第一,数据质量,而不是组装器本身,对组装基因组的质量有巨大影响。第二,组装的连续程度在不同组装器和不同基因组之间存在巨大差异;第三,程序集的正确性也有很大差异,并且与连续性统计数据没有很好的相关性。为了使其他人能够复制我们的结果,我们的所有数据和方法以及本研究中使用的所有组装程序都是免费提供的。
New sequencing technology has dramatically altered the landscape of whole-genome sequencing, allowing scientists to initiate numerous projects to decode the genomes of previously unsequenced organisms. The lowest-cost technology can generate deep coverage of most species, including mammals, in just a few days. The sequence data generated by one of these projects consist of millions or billions of short DNA sequences (reads) that range from 50 to 150 it in length. These sequences must then be assembled de novo before most genome analyses can begin. Unfortunately, genome assembly remains a very difficult problem, made more difficult by shorter reads and unreliable long-range linking information. In this study, we evaluated several of the leading de novo assembly algorithms on four different short-read data sets, all generated by Illumina sequencers. Our results describe the relative performance of the different assemblers as well as other significant differences in assembly difficulty that appear to be inherent in the genomes themselves. Three over-arching conclusions are apparent: first, that data quality, rather than the assembler itself, has a dramatic effect on the quality of an assembled genome., second, that the degree of contiguity of an assembly varies enormously among different assemblers and different genomes; and third, that the correctness of an assembly also varies widely and is not well correlated with statistics on contiguity. To enable others to replicate our results, all of our data and methods are freely available, as are all assemblers used in this study.