Detrimental effects of duplicate reads and low complexity regions on RNA- and ChIP-seq data.

Detrimental effects of duplicate reads and low complexity regions on RNA- and ChIP-seq data.
复制标题

DOI:
10.1186/1471-2105-16-s13-s10
复制
发表时间:
2015
期刊:
影响因子:
3
通讯作者:
Wren JD
Wren JD
中科院分区:
生物学4区
文献类型:
--
作者:
Dozmorov MG;Adrianto I;Giles CB;Glass E;Glenn SB;Montgomery C;Sivils KL;Olson LE;Iwayama T;Freeman WM;Lessard CJ;Wren JD

文献摘要

被引文献

相似文献

适配器修剪和去除重复读数在下一代测序管道中是常见的做法。测序读段含糊地映射到重复和低复杂性区域,也可能对准确评估生物信号造成问题,但它们对测序数据的影响尚未得到太多关注。在RNA和ChIP-seq实验中,我们研究了修剪适配器、去除重复和过滤掉重叠低复杂性区域的reads如何影响生物信号的重要性。我们评估了数据处理步骤对RNA-和ChIP-seq数据的比对统计和功能富集分析结果的影响。我们比较了相同患者样本上差异处理的RNA-seq数据与匹配的微阵列数据,以确定预处理的变化是否改善了两者之间的相关性。我们开发了一个简单的工具来去除低复杂性区域,RepeatSoaker,可在https://github.com/mdozmorov/RepeatSoaker上获得,并测试了它对比对统计和富集分析结果的影响。适配器修剪和重复去除都适度提高了RNA-seq和ChIP-seq数据中的生物信号强度。对RepeatMasker定义的低复杂度区域重叠的reads进行积极过滤,进一步提高了生物信号的强度,以及RNA-seq与微阵列基因表达数据之间的相关性。在RNA-seq和ChIP-seq数据中,适配器修剪和重复去除,加上过滤掉重叠低复杂度区域的reads,可以提高检测生物信号的质量和可靠性。
Adapter trimming and removal of duplicate reads are common practices in next-generation sequencing pipelines. Sequencing reads ambiguously mapped to repetitive and low complexity regions can also be problematic for accurate assessment of the biological signal, yet their impact on sequencing data has not received much attention. We investigate how trimming the adapters, removing duplicates, and filtering out reads overlapping low complexity regions influence the significance of biological signal in RNA- and ChIP-seq experiments. We assessed the effect of data processing steps on the alignment statistics and the functional enrichment analysis results of RNA- and ChIP-seq data. We compared differentially processed RNA-seq data with matching microarray data on the same patient samples to determine whether changes in pre-processing improved correlation between the two. We have developed a simple tool to remove low complexity regions, RepeatSoaker, available at https://github.com/mdozmorov/RepeatSoaker, and tested its effect on the alignment statistics and the results of the enrichment analyses. Both adapter trimming and duplicate removal moderately improved the strength of biological signals in RNA-seq and ChIP-seq data. Aggressive filtering of reads overlapping with low complexity regions, as defined by RepeatMasker, further improved the strength of biological signals, and the correlation between RNA-seq and microarray gene expression data. Adapter trimming and duplicates removal, coupled with filtering out reads overlapping low complexity regions, is shown to increase the quality and reliability of detecting biological signals in RNA-seq and ChIP-seq data.