Indel-correcting DNA barcodes for high-throughput sequencing.
Indel-correcting DNA barcodes for high-throughput sequencing.
复制标题
DOI:
10.1073/pnas.1802640115
复制
发表时间:
2018-07-03
影响因子:
11.1
通讯作者:
Press WH
中科院分区:
文献类型:
--
作者:
Hawkins JA;Jones SK Jr;Finkelstein IJ;Press WH
Modern high-throughput biological assays study pooled populations of individual members by labeling each member with a unique DNA sequence called a “barcode.” DNA barcodes are frequently corrupted by DNA synthesis and sequencing errors, leading to significant data loss and incorrect data interpretation. Here, we describe an error correction strategy to improve the efficiency and statistical power of DNA barcodes. Our strategy accurately handles insertions and deletions (indels) in DNA barcodes, the most common type of error encountered during DNA synthesis and sequencing, resulting in order-of-magnitude increases in accuracy, efficiency, and signal-to-noise ratio. The accompanying software package makes deployment of these barcodes straightforward for the broader experimental scientist community. Many large-scale, high-throughput experiments use DNA barcodes, short DNA sequences prepended to DNA libraries, for identification of individuals in pooled biomolecule populations. However, DNA synthesis and sequencing errors confound the correct interpretation of observed barcodes and can lead to significant data loss or spurious results. Widely used error-correcting codes borrowed from computer science (e.g., Hamming, Levenshtein codes) do not properly account for insertions and deletions (indels) in DNA barcodes, even though deletions are the most common type of synthesis error. Here, we present and experimentally validate filled/truncated right end edit (FREE) barcodes, which correct substitution, insertion, and deletion errors, even when these errors alter the barcode length. FREE barcodes are designed with experimental considerations in mind, including balanced guanine-cytosine (GC) content, minimal homopolymer runs, and reduced internal hairpin propensity. We generate and include lists of barcodes with different lengths and error correction levels that may be useful in diverse high-throughput applications, including >106 single-error–correcting 16-mers that strike a balance between decoding accuracy, barcode length, and library size. Moreover, concatenating two or more FREE codes into a single barcode increases the available barcode space combinatorially, generating lists with >1015 error-correcting barcodes. The included software for creating barcode libraries and decoding sequenced barcodes is efficient and designed to be user-friendly for the general biology community.
登录
查看更多内容
影响因子:
3
作者:
Buschmann T;Bystrykh LV
通讯作者:
Bystrykh LV
影响因子:
46.9
作者:
Kitzman, Jacob O.
通讯作者:
Kitzman, Jacob O.
影响因子:
64.5
作者:
Macosko EZ;Basu A;Satija R;Nemesh J;Shekhar K;Goldman M;Tirosh I;Bialas AR;Kamitaki N;Martersteck EM;Trombetta JJ;Weitz DA;Sanes JR;Shalek AK;Regev A;McCarroll SA
通讯作者:
McCarroll SA
影响因子:
14.9
作者:
Lee DF;Lu J;Chang S;Loparo JJ;Xie XS
通讯作者:
Xie XS
影响因子:
14.8
作者:
Zilionis, Rapolas;Nainys, Juozas;Mazutis, Linas
通讯作者:
Mazutis, Linas