Levenshtein error-correcting barcodes for multiplexed DNA sequencing.

Levenshtein error-correcting barcodes for multiplexed DNA sequencing.
复制标题

DOI:
10.1186/1471-2105-14-272
复制
发表时间:
2013-09-11
期刊:
影响因子:
3
通讯作者:
Bystrykh LV
Bystrykh LV
中科院分区:
生物学4区
文献类型:
--
作者:
Buschmann T;Bystrykh LV

文献摘要

参考文献

被引文献

相似文献

高通量测序技术在质量、能力和成本方面不断改进,为DNA和RNA研究提供了多方面的应用。对于小基因组或较大基因组的一部分,DNA样品可以混合并一起加载到相同的测序轨道上。这种所谓的多重方法依赖于特定的DNA标签或条形码,其连接到测序或扩增引物上,因此出现在每次读取的序列开始处。测序后,基于各自的条形码序列鉴定每个样品读段。在合成、引物连接、DNA扩增或测序期间DNA条形码的改变可能导致不正确的样品鉴定,除非错误被揭示并纠正。这可以通过实现纠错算法和代码来实现。这种条形码化策略增加了正确鉴定的样品的总数,从而提高了总体测序效率。两种流行的纠错码是汉明码和Levenshtein码。Levenshtein码只对已知长度的字起作用。由于具有嵌入式条形码的DNA序列基本上是一个连续的长字,因此经典Levenshtein算法的应用是有问题的。在本文中,我们证明了降低的错误校正能力的Levenshtein代码在DNA的上下文中,并建议适应Levenshtein代码,被证明有效地纠正DNA序列中的核苷酸错误。在我们的改编中,我们考虑到DNA上下文,并在插入或删除被发现时重新定义单词长度。在模拟中,我们展示了优越的上级纠错能力的新方法相比,传统的Levenshtein和汉明为基础的代码中存在的多个错误。我们提出了一个适应Levenshtein代码的DNA背景下能够校正的预定义数量的插入,缺失和取代突变。我们的改进方法是另外能够恢复的新长度的损坏的码字和纠正平均更多的随机突变比传统的Levenshtein或汉明码。作为这项工作的一部分,我们准备了软件,用于基于我们的新方法灵活生成DNA代码。为了使代码适应特定的实验条件,用户可以自定义序列过滤、可纠正突变的数量和条形码长度,以获得最高性能。
High-throughput sequencing technologies are improving in quality, capacity and costs, providing versatile applications in DNA and RNA research. For small genomes or fraction of larger genomes, DNA samples can be mixed and loaded together on the same sequencing track. This so-called multiplexing approach relies on a specific DNA tag or barcode that is attached to the sequencing or amplification primer and hence appears at the beginning of the sequence in every read. After sequencing, each sample read is identified on the basis of the respective barcode sequence. Alterations of DNA barcodes during synthesis, primer ligation, DNA amplification, or sequencing may lead to incorrect sample identification unless the error is revealed and corrected. This can be accomplished by implementing error correcting algorithms and codes. This barcoding strategy increases the total number of correctly identified samples, thus improving overall sequencing efficiency. Two popular sets of error-correcting codes are Hamming codes and Levenshtein codes. Levenshtein codes operate only on words of known length. Since a DNA sequence with an embedded barcode is essentially one continuous long word, application of the classical Levenshtein algorithm is problematic. In this paper we demonstrate the decreased error correction capability of Levenshtein codes in a DNA context and suggest an adaptation of Levenshtein codes that is proven of efficiently correcting nucleotide errors in DNA sequences. In our adaption we take the DNA context into account and redefine the word length whenever an insertion or deletion is revealed. In simulations we show the superior error correction capability of the new method compared to traditional Levenshtein and Hamming based codes in the presence of multiple errors. We present an adaptation of Levenshtein codes to DNA contexts capable of correction of a pre-defined number of insertion, deletion, and substitution mutations. Our improved method is additionally capable of recovering the new length of the corrupted codeword and of correcting on average more random mutations than traditional Levenshtein or Hamming codes. As part of this work we prepared software for the flexible generation of DNA codes based on our new approach. To adapt codes to specific experimental conditions, the user can customize sequence filtering, the number of correctable mutations and barcode length for highest performance.
DOI: 10.1371/journal.pone.0036852
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Bystrykh LV
通讯作者: Bystrykh LV
DOI: 10.1038/nprot.2009.64
发表时间: 2009
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Uren, Anthony G.;Mikkers, Harald;Kool, Jaap;van der Weyden, Louise;Lund, Anders H.;Wilson, Catherine H.;Rance, Richard;Jonkers, Jos;van Lohuizen, Maarten;Berns, Anton;Adams, David J.
通讯作者: Adams, David J.
DOI: 10.1023/a:1011275112159
发表时间: 2001-01-01
影响因子: 1.6
作者:
Bogdanova, GT;Brouwer, AE;Östergård, PRJ
通讯作者: Östergård, PRJ
DOI: 10.1093/nar/gkm760
发表时间: 2007
影响因子: 14.9
作者:
Parameswaran P;Jalili R;Tao L;Shokralla S;Gharizadeh B;Ronaghi M;Fire AZ
通讯作者: Fire AZ
DOI: 10.1186/1471-2164-11-716
发表时间: 2010-12-20
期刊: BMC genomics
影响因子: 4.4
作者:
Buermans HP;Ariyurek Y;van Ommen G;den Dunnen JT;'t Hoen PA
通讯作者: 't Hoen PA