Syotti: scalable bait design for DNA enrichment.

Syotti: scalable bait design for DNA enrichment.
复制标题

DOI:
10.1093/bioinformatics/btac226
复制
发表时间:
2022-06-24
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

诱饵富集是一种越来越普遍的方案,因为它已被证明可以成功扩增宏基因组样品中的感兴趣区域。在该方法中,设计、制造一组合成探针(“诱饵”)并将其应用于片段化的宏基因组DNA。探针与片段化的DNA结合,并且冲洗掉任何未结合的DNA,留下待扩增的结合片段用于测序。Metsky等人证明,诱饵富集能够检测宏基因组样品中的大量人类病毒病原体。我们通过定义最小诱饵覆盖问题来正式化设计诱饵的问题,表明即使在非常严格的假设下,该问题也是NP难的,并设计了一个有效的启发式算法,利用简洁的数据结构。我们将我们的方法称为Syotti。Syotti的运行时间在实践中显示出线性缩放,比最先进的方法(包括Metsky等人的方法)至少快一个数量级。同时,我们的方法产生的诱饵集小于竞争方法产生的诱饵集,同时也留下更少的未覆盖位置。最后,我们表明,Syotti只需要25分钟就可以设计出一个由来自1000个相关细菌亚株的30亿个核苷酸组成的数据集的诱饵,而Metsky等人的方法显示出明显的超线性运行时间,并且在72小时内甚至不能处理17%的数据子集。 https://github.com/jnalanko/syotti. 补充数据可在Bioinformatics在线获得。
Bait enrichment is a protocol that is becoming increasingly ubiquitous as it has been shown to successfully amplify regions of interest in metagenomic samples. In this method, a set of synthetic probes (‘baits’) are designed, manufactured and applied to fragmented metagenomic DNA. The probes bind to the fragmented DNA and any unbound DNA is rinsed away, leaving the bound fragments to be amplified for sequencing. Metsky et al. demonstrated that bait-enrichment is capable of detecting a large number of human viral pathogens within metagenomic samples. We formalize the problem of designing baits by defining the Minimum Bait Cover problem, show that the problem is NP-hard even under very restrictive assumptions, and design an efficient heuristic that takes advantage of succinct data structures. We refer to our method as Syotti. The running time of Syotti shows linear scaling in practice, running at least an order of magnitude faster than state-of-the-art methods, including the method of Metsky et al. At the same time, our method produces bait sets that are smaller than the ones produced by the competing methods, while also leaving fewer positions uncovered. Lastly, we show that Syotti requires only 25 min to design baits for a dataset comprised of 3 billion nucleotides from 1000 related bacterial substrains, whereas the method of Metsky et al. shows clearly super-linear running time and fails to process even a subset of 17% of the data in 72 h. https://github.com/jnalanko/syotti. Supplementary data are available at Bioinformatics online.
DOI: 10.1038/s41587-018-0006-x
发表时间: 2019-02-01
影响因子: 46.9
作者:
Metsky, Hayden C.;Siddle, Katherine J.;Matranga, Christian B.
通讯作者: Matranga, Christian B.
DOI: 10.1145/1082036.1082039
发表时间: 2005-07-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Ferragina, P;Manzini, G
通讯作者: Manzini, G
DOI: 10.1128/aac.01324-19
发表时间: 2020-01-01
影响因子: 4.9
作者:
Guitor, Allison K.;Raphenya, Amogelang R.;Wright, Gerard D.
通讯作者: Wright, Gerard D.
DOI: 10.7717/peerj.2584
发表时间: 2016-10-18
期刊: PEERJ
影响因子: 2.7
作者:
Rognes, Torbjorn;Flouri, Tomas;Mahe, Frederic
通讯作者: Mahe, Frederic
DOI: 10.1186/s40168-017-0361-8
发表时间: 2017-10-17
期刊: Microbiome
影响因子: 15.5
作者:
Noyes NR;Weinroth ME;Parker JK;Dean CJ;Lakin SM;Raymond RA;Rovira P;Doster E;Abdo Z;Martin JN;Jones KL;Ruiz J;Boucher CA;Belk KE;Morley PS
通讯作者: Morley PS