Minerva: an alignment- and reference-free approach to deconvolve Linked-Reads for metagenomics

Minerva: an alignment- and reference-free approach to deconvolve Linked-Reads for metagenomics
复制标题

DOI:
10.1101/gr.235499.118
复制
发表时间:
2019-01-01
期刊:
影响因子:
7
通讯作者:
Hajirasouliha, Iman
Hajirasouliha, Iman
中科院分区:
生物学1区
文献类型:
--
作者:
Danko, David C.;Meleshko, Dmitry;Hajirasouliha, Iman

文献摘要

被引文献

相似文献

新兴的Linked-Read技术(又称读云或条形码短读)重新唤起了人们对短读技术的兴趣,将其作为了解基因组和宏基因组中大规模结构的可行方法。Linked-Read技术,如10 x Chromium系统,使用微流体系统和一组专门的3'条形码(又名UlD)来标记来源于相同长DNA片段的短DNA读数;随后,在标准短读数平台上对标记的读数进行测序。这种方法导致了有趣的妥协。DNA的每个长片段仅稀疏地被读段覆盖,没有关于来自相同片段的读段的排序的信息被保留,并且3'条形码匹配来自大约2-20个长DNA片段的读段。然而,与长读段技术相比,测序的每个碱基的成本要低得多,需要的输入DNA要少得多,并且每个碱基的错误率与Illumina短读段相同。在本文中,我们正式描述了一个特定的算法问题,共同的Linked-Read技术:一个单一的3'条形码到集群,代表单一的长片段的DNA的读解卷积。我们介绍Minerva,一种基于图形的算法,近似解决了宏基因组数据的条形码去卷积问题(其中参考基因组可能不完整或不可用)。此外,我们开发了两个示范,其中条形码读取的去卷积改善了下游结果,提高了分类分配和基于k聚体的聚类的特异性。据我们所知,我们是第一个解决宏基因组学中条形码去卷积问题的人。
Emerging Linked-Read technologies (aka read cloud or barcoded short-reads) have revived interest in short-read technology as a viable approach to understand large-scale structures in genomes and metagenomes. Linked-Read technologies, such as the 10x Chromium system, use a microfluidic system and a specialized set of 3' barcodes (aka UlDs) to tag short DNA reads sourced from the same long fragment of DNA; subsequently, the tagged reads are sequenced on standard short-read platforms. This approach results in interesting compromises. Each long fragment of DNA is only sparsely covered by reads, no information about the ordering of reads from the same fragment is preserved, and 3' barcodes match reads from roughly 2-20 long fragments of DNA. However, compared to long-read technologies, the cost per base to sequence is far lower, far less input DNA is required, and the per base error rate is that of Illumina short-reads. In this paper, we formally describe a particular algorithmic issue common to Linked-Read technology: the deconvolution of reads with a single 3' barcode into clusters that represent single long fragments of DNA. We introduce Minerva, a graph-based algorithm that approximately solves the barcode deconvolution problem for metagenomic data (where reference genomes may be incomplete or unavailable). Additionally, we develop two demonstrations where the deconvolution of barcoded reads improves downstream results, improving the specificity of taxonomic assignments and of k-mer-based clustering. To the best of our knowledge, we are the first to address the problem of barcode deconvolution in metagenomics.