Apollo: a sequencing-technology-independent, scalable and accurate assembly polishing algorithm

Apollo: a sequencing-technology-independent, scalable and accurate assembly polishing algorithm
复制标题

DOI:
10.1093/bioinformatics/btaa179
复制
发表时间:
2020-06-15
期刊:
影响因子:
5.8
通讯作者:
Mutlu, Onur
Mutlu, Onur
中科院分区:
生物学3区
文献类型:
--
作者:
Firtina, Can;Kim, Jeremie S.;Mutlu, Onur

文献摘要

被引文献

相似文献

动机:第三代测序技术可以测序包含多达200万个碱基对的长读段。这些长读段用于构建组装体(即受试者的基因组),其进一步用于下游基因组分析。不幸的是,第三代测序技术具有高测序错误率,并且这些长读段中的大部分碱基对被错误地识别。这些错误传播到组件并影响基因组分析的准确性。组装抛光算法通过使用来自读段和组装之间的比对的信息(即读段到组装比对信息)抛光或修复组装中的错误来最小化这种错误传播。然而,目前的组装抛光算法只能使用来自某种测序技术或小组装的读数来抛光组装。这种技术依赖性和组装大小依赖性要求研究人员(i)运行多个抛光算法和(ii)使用大基因组的小块来分别使用所有可用的读集和抛光大基因组。我们介绍阿波罗,一个通用的装配抛光算法,可以很好地扩展以抛光任何尺寸的装配使用来自所有测序技术(即第二代和第三代)的读数,可以对所有的基因组(即大基因组和小基因组)进行测序。我们的目标是提供一种单一的算法,该算法使用来自所有可用测序技术的读取集来提高组装抛光的准确性,并且可以抛光大基因组。Apollo(i)将组装件建模为简档隐马尔可夫模型(pHMM),(ii)使用读取到组装件比对来用前向-后向算法训练pHMM,以及(iii)用维特比算法解码经训练的模型以产生抛光组装件。我们对真实的读段集的实验表明,Apollo是唯一的算法,它(i)在单次运行中使用来自任何测序技术的读段,(ii)可以很好地扩展以抛光大型组装件,而不会将组装件分成多个部分。
Motivation: Third-generation sequencing technologies can sequence long reads that contain as many as 2 million base pairs. These long reads are used to construct an assembly (i.e. the subject's genome), which is further used in downstream genome analysis. Unfortunately, third-generation sequencing technologies have high sequencing error rates and a large proportion of base pairs in these long reads is incorrectly identified. These errors propagate to the assembly and affect the accuracy of genome analysis. Assembly polishing algorithms minimize such error propagation by polishing or fixing errors in the assembly by using information from alignments between reads and the assembly (i.e. read-to-assembly alignment information). However, current assembly polishing algorithms can only polish an assembly using reads from either a certain sequencing technology or a small assembly. Such technology-dependency and assembly-size dependency require researchers to (i) run multiple polishing algorithms and (ii) use small chunks of a large genome to use all available readsets and polish large genomes, respectively.Results: We introduce Apollo, a universal assembly polishing algorithm that scales well to polish an assembly of any size (i.e. both large and small genomes) using reads from all sequencing technologies (i.e. second- and third generation). Our goal is to provide a single algorithm that uses read sets from all available sequencing technologies to improve the accuracy of assembly polishing and that can polish large genomes. Apollo (i) models an assembly as a profile hidden Markov model (pHMM), (ii) uses read-to-assembly alignment to train the pHMM with the Forward-Backward algorithm and (iii) decodes the trained model with the Viterbi algorithm to produce a polished assembly. Our experiments with real readsets demonstrate that Apollo is the only algorithm that (i) uses reads from any sequencing technology within a single run and (ii) scales well to polish large assemblies without splitting the assembly into multiple parts.