Detection and removal of biases in the analysis of next-generation sequencing reads.

Detection and removal of biases in the analysis of next-generation sequencing reads.
复制标题

DOI:
10.1371/journal.pone.0016685
复制
发表时间:
2011-01-31
期刊:
影响因子:
3.7
通讯作者:
Ast G
Ast G
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Schwartz S;Oren R;Ast G

文献摘要

参考文献

被引文献

相似文献

自从下一代测序(NGS)技术出现以来,已经投入了大量精力来开发用于分析短读段的工具。与此同时,人们对这些技术中固有的偏见的了解也在增加。在这里,我们讨论了我们在分析各种Illumina数据集时遇到的四种不同的偏差。这些偏差是由于生物学和统计学效应,特别是影响不同基因组区域之间的比较。具体而言,我们遇到了与核苷酸在测序循环中的分布、可映射性、前mRNA与mRNA的污染以及RNA的不均匀水解有关的偏差。这些偏差中的大多数并不特定于一个分析的数据集,而是存在于各种数据集和各种基因组背景中。重要的是,这些偏差中的一些与生物学特征(包括转录本长度、基因表达水平、保守水平和外显子-内含子结构)高度显著相关,从而误导性地增加了结果的可信度。我们还证明了这些偏见的相关性,在分析的NGS数据集映射转录参与RNA聚合酶II(RNAPII)的背景下,外显子-内含子架构,并表明,消除这些偏见是至关重要的,以避免错误的数据解释。总的来说,我们的研究结果突出了NGS读取分析中的几个重要陷阱,挑战和方法。
Since the emergence of next-generation sequencing (NGS) technologies, great effort has been put into the development of tools for analysis of the short reads. In parallel, knowledge is increasing regarding biases inherent in these technologies. Here we discuss four different biases we encountered while analyzing various Illumina datasets. These biases are due to both biological and statistical effects that in particular affect comparisons between different genomic regions. Specifically, we encountered biases pertaining to the distributions of nucleotides across sequencing cycles, to mappability, to contamination of pre-mRNA with mRNA, and to non-uniform hydrolysis of RNA. Most of these biases are not specific to one analyzed dataset, but are present across a variety of datasets and within a variety of genomic contexts. Importantly, some of these biases correlated in a highly significant manner with biological features, including transcript length, gene expression levels, conservation levels, and exon-intron architecture, misleadingly increasing the credibility of results due to them. We also demonstrate the relevance of these biases in the context of analyzing an NGS dataset mapping transcriptionally engaged RNA polymerase II (RNAPII) in the context of exon-intron architecture, and show that elimination of these biases is crucial for avoiding erroneous interpretation of the data. Collectively, our results highlight several important pitfalls, challenges and approaches in the analysis of NGS reads.
DOI: 10.1186/gb-2009-10-8-r83
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Kircher M;Stenzel U;Kelso J
通讯作者: Kelso J
DOI: 10.1093/nar/gkh023
发表时间: 2004-01-01
影响因子: 14.9
作者:
Griffiths-Jones, S
通讯作者: Griffiths-Jones, S
DOI: 10.1038/ng.322
发表时间: 2009-03
期刊: Nature genetics
影响因子: 30.8
作者:
Kolasinska-Zwierz P;Down T;Latorre I;Liu T;Liu XS;Ahringer J
通讯作者: Ahringer J
DOI: 10.1093/nar/gkj112
发表时间: 2006-01-01
影响因子: 14.9
作者:
Griffiths-Jones S;Grocock RJ;van Dongen S;Bateman A;Enright AJ
通讯作者: Enright AJ
DOI: 10.1371/journal.pcbi.1000566
发表时间: 2009-11
影响因子: 4.3
作者:
Hon G;Wang W;Ren B
通讯作者: Ren B