A fast adaptive algorithm for computing whole-genome homology maps

A fast adaptive algorithm for computing whole-genome homology maps
复制标题

DOI:
10.1101/259986
复制
发表时间:
2018-02
期刊:
影响因子:
5.8
通讯作者:
Chirag Jain;S. Koren;A. Dilthey;A. Phillippy;S. Aluru
Chirag Jain;S. Koren;A. Dilthey;A. Phillippy;S. Aluru
中科院分区:
生物学3区
文献类型:
--
作者:
Chirag Jain;S. Koren;A. Dilthey;A. Phillippy;S. Aluru

文献摘要

被引文献

相似文献

在基因组学中,全基因组比对是一个重要的问题,用于比较不同物种,绘制草图组装到参考基因组,以及识别重复序列。然而,对于大型植物和动物基因组,这项任务仍然是计算和内存密集型的。结果介绍了一种计算长DNA序列局部比对边界的近似算法。给定最小对齐长度和标识阈值,我们的算法使用基于kmer的统计计算所需的对齐边界和标识估计,并在输出灵敏度上保持足够的概率保证。此外,为了优先考虑更高的评分对齐间隔,我们开发了一种基于平面扫描的滤波技术,该技术在理论上是最优的,在实践中是有效的。这些想法的实施导致了一个快速和准确的组装到基因组和基因组到基因组的制图器。结果,我们能够在大约1分钟的总执行时间内将错误纠正的NA12878全基因组序列映射到hg38人类参考基因组,并且在多个数据集上达到97%。最后,我们对人类基因组进行了敏感的自比对,以计算长度≥1 Kbp和≥90%同源性的所有重复。报告的输出达到了很好的召回率,并且比当前UCSC基因组浏览器的片段重复注释多覆盖了5%的碱基。可用性https://github.com/marbl/MashMap联系adam.phillippy@nih.gov, aluru@cc.gatech.edu
Motivation Whole-genome alignment is an important problem in genomics for comparing different species, mapping draft assemblies to reference genomes, and identifying repeats. However, for large plant and animal genomes, this task remains compute and memory intensive. Results We introduce an approximate algorithm for computing local alignment boundaries between long DNA sequences. Given a minimum alignment length and an identity threshold, our algorithm computes the desired alignment boundaries and identity estimates using kmer-based statistics, and maintains sufficient probabilistic guarantees on the output sensitivity. Further, to prioritize higher scoring alignment intervals, we develop a plane-sweep based filtering technique which is theoretically optimal and practically efficient. Implementation of these ideas resulted in a fast and accurate assembly-to-genome and genome-to-genome mapper. As a result, we were able to map an error-corrected whole-genome NA12878 human assembly to the hg38 human reference genome in about one minute total execution time and 97% on multiple datasets. Finally, we performed a sensitive self-alignment of the human genome to compute all duplications of length ≥ 1 Kbp and ≥ 90% identity. The reported output achieves good recall and covers 5% more bases than the current UCSC genome browser’s segmental duplication annotation. Availability https://github.com/marbl/MashMap Contact adam.phillippy@nih.gov, aluru@cc.gatech.edu