Assembly of long, error-prone reads using repeat graphs

Assembly of long, error-prone reads using repeat graphs
复制标题

DOI:
10.1038/s41587-019-0072-8
复制
发表时间:
2019-05-01
影响因子:
46.9
通讯作者:
Pevzner, Pavel A.
Pevzner, Pavel A.
中科院分区:
工程技术1区
文献类型:
--
作者:
Kolmogorov, Mikhail;Yuan, Jeffrey;Pevzner, Pavel A.

文献摘要

被引文献

相似文献

精确的基因组组装受到重复区域的阻碍。虽然长的单分子测序读数比短读数数据更能解析基因组重复,但大多数长读数组装算法不能提供产生最佳组装所需的重复特征。在这里,我们提出了Flye,一个长读汇编算法,它在一个未知的重复图中生成任意路径,称为不相交,并从这些错误百出的不相交构造一个准确的重复图。我们将Flye与五个最先进的装配机进行比较,结果表明,它生成的装配量更好或相当,同时速度快了一个数量级。与现有的组装程序相比,Flye几乎将人类基因组组装的邻接性(以NGA50组装质量指标衡量)增加了一倍。
Accurate genome assembly is hampered by repetitive regions. Although long single molecule sequencing reads are better able to resolve genomic repeats than short-read data, most long-read assembly algorithms do not provide the repeat characterization necessary for producing optimal assemblies. Here, we present Flye, a long-read assembly algorithm that generates arbitrary paths in an unknown repeat graph, called disjointigs, and constructs an accurate repeat graph from these error-riddled disjointigs. We benchmark Flye against five state-of-the-art assemblers and show that it generates better or comparable assemblies, while being an order of magnitude faster. Flye nearly doubled the contiguity of the human genome assembly (as measured by the NGA50 assembly quality metric) compared with existing assemblers.