Elimination of PCR duplicates in RNA-seq and small RNA-seq using unique molecular identifiers.

Elimination of PCR duplicates in RNA-seq and small RNA-seq using unique molecular identifiers.
复制标题

DOI:
10.1186/s12864-018-4933-1
复制
发表时间:
2018-07-13
期刊:
影响因子:
4.4
通讯作者:
Weng Z
Weng Z
中科院分区:
生物学2区
文献类型:
--
作者:
Fu Y;Wu PH;Beane T;Zamore PD;Weng Z

文献摘要

参考文献

被引文献

相似文献

RNA-seq和small RNA-seq是研究基因调控和功能的强大定量工具。常见的高通量测序方法依赖于聚合酶链式反应(PCR)来扩增起始材料,但并非每个分子都能同等地扩增,导致一些分子被过度代表。独特分子标识符(UMI)可用于区分源自单个分子的不期望的PCR重复和来自不同分子的相同但有生物学意义的读段。我们已经将UMI整合到RNA-seq和小型RNA-seq协议中,并开发了分析所得数据的工具。我们的UMI包含随机核苷酸的延伸,其长度足以捕获从小鼠睾丸产生的RNA-seq和小RNA-seq文库中的不同分子种类。我们的方法产生高质量的数据,同时允许在高深度库中对所有分子进行独特标记。使用模拟和真实的数据集,我们证明了我们的方法提高了RNA-seq和小RNA-seq数据的再现性。值得注意的是,我们发现起始材料的量和测序深度,而不是PCR循环的数量,决定PCR重复频率。最后,我们表明,计算删除PCR重复的基础上,只有他们的映射坐标引入大量的数据分析偏差。本文的在线版本(10.1186/s12864-018-4933-1)包含补充材料,可供授权用户使用。
RNA-seq and small RNA-seq are powerful, quantitative tools to study gene regulation and function. Common high-throughput sequencing methods rely on polymerase chain reaction (PCR) to expand the starting material, but not every molecule amplifies equally, causing some to be overrepresented. Unique molecular identifiers (UMIs) can be used to distinguish undesirable PCR duplicates derived from a single molecule and identical but biologically meaningful reads from different molecules. We have incorporated UMIs into RNA-seq and small RNA-seq protocols and developed tools to analyze the resulting data. Our UMIs contain stretches of random nucleotides whose lengths sufficiently capture diverse molecule species in both RNA-seq and small RNA-seq libraries generated from mouse testis. Our approach yields high-quality data while allowing unique tagging of all molecules in high-depth libraries. Using simulated and real datasets, we demonstrate that our methods increase the reproducibility of RNA-seq and small RNA-seq data. Notably, we find that the amount of starting material and sequencing depth, but not the number of PCR cycles, determine PCR duplicate frequency. Finally, we show that computational removal of PCR duplicates based only on their mapping coordinates introduces substantial bias into data analysis. The online version of this article (10.1186/s12864-018-4933-1) contains supplementary material, which is available to authorized users.
DOI: 10.1186/s13059-015-0684-3
发表时间: 2015-06-06
期刊: Genome biology
影响因子: 12.3
作者:
Bose S;Wan Z;Carr A;Rizvi AH;Vieira G;Pe'er D;Sims PA
通讯作者: Sims PA
DOI: 10.1093/bioinformatics/btu647
发表时间: 2015-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Han BW;Wang W;Zamore PD;Weng Z
通讯作者: Weng Z
DOI: 10.1186/s12864-015-1788-6
发表时间: 2015-08-05
期刊: BMC GENOMICS
影响因子: 4.4
作者:
Collins, John E.;Wali, Neha;Busch-Nentwich, Elisabeth M.
通讯作者: Busch-Nentwich, Elisabeth M.
DOI: 10.1021/ac500459p
发表时间: 2014-03-18
影响因子: 7.4
作者:
Fu, Glenn K.;Wilhelmy, Julie;Stern, David;Fan, H. Christina;Fodor, Stephen P. A.
通讯作者: Fodor, Stephen P. A.
DOI: 10.1016/j.cub.2008.04.042
发表时间: 2008-05-20
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Addo-Quaye, Charles;Eshoo, Tifani W.;Axtell, Michael J.
通讯作者: Axtell, Michael J.