Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the MaSuRCA mega-reads algorithm.

Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the MaSuRCA mega-reads algorithm.
复制标题

DOI:
10.1101/gr.213405.116
复制
发表时间:
2017-05
期刊:
影响因子:
7
通讯作者:
Salzberg SL
Salzberg SL
中科院分区:
生物学1区
文献类型:
--
作者:
Zimin AV;Puiu D;Luo MC;Zhu T;Koren S;Marçais G;Yorke JA;Dvořák J;Salzberg SL

文献摘要

被引文献

相似文献

通过单分子测序技术产生的长测序读段提供了显著改善基因组组装的邻接性的可能性。目前最大的挑战是长时间读取具有相对较高的错误率,目前约为15%。高错误率使得很难单独使用这些数据,特别是对于高度重复的植物基因组。原始数据中的错误可能导致共有基因组序列中的插入或缺失错误(indel),这反过来又为下游分析造成重大问题;例如,单个indel可能移动阅读框并错误地截短蛋白质序列。在这里,我们描述了一种算法,该算法通过将长的高错误读段与较短但更准确的Illumina测序读段相结合来解决高错误率问题,其错误率平均<1%。我们的混合组装算法结合了这两种类型的读段来构建长而准确的mega-reads,然后使用CABOG组装器组装mega-reads,该组装器是为长读段设计的。我们将这种技术应用于来自节节麦物种的Illumina和PacBio序列的大型数据集,节节麦是一种大型且极其重复的植物基因组,其抵抗了先前的组装尝试。我们表明,所得组装的重叠群远远大于任何以前的组装,N50重叠群大小为486,807个核苷酸。我们将重叠群与独立产生的光学图谱进行比较,以评估其大规模的准确性,并与一组高质量的基于细菌人工染色体(BAC)的组件进行比较,以评估基本水平的准确性。
Long sequencing reads generated by single-molecule sequencing technology offer the possibility of dramatically improving the contiguity of genome assemblies. The biggest challenge today is that long reads have relatively high error rates, currently around 15%. The high error rates make it difficult to use this data alone, particularly with highly repetitive plant genomes. Errors in the raw data can lead to insertion or deletion errors (indels) in the consensus genome sequence, which in turn create significant problems for downstream analysis; for example, a single indel may shift the reading frame and incorrectly truncate a protein sequence. Here, we describe an algorithm that solves the high error rate problem by combining long, high-error reads with shorter but much more accurate Illumina sequencing reads, whose error rates average <1%. Our hybrid assembly algorithm combines these two types of reads to construct mega-reads, which are both long and accurate, and then assembles the mega-reads using the CABOG assembler, which was designed for long reads. We apply this technique to a large data set of Illumina and PacBio sequences from the species Aegilops tauschii, a large and extremely repetitive plant genome that has resisted previous attempts at assembly. We show that the resulting assembled contigs are far larger than in any previous assembly, with an N50 contig size of 486,807 nucleotides. We compare the contigs to independently produced optical maps to evaluate their large-scale accuracy, and to a set of high-quality bacterial artificial chromosome (BAC)-based assemblies to evaluate base-level accuracy.