Reducing bias in RNA sequencing data: a novel approach to compute counts.

Reducing bias in RNA sequencing data: a novel approach to compute counts.
复制标题

DOI:
10.1186/1471-2105-15-s1-s7
复制
发表时间:
2014
期刊:
影响因子:
3
通讯作者:
Di Camillo B
Di Camillo B
中科院分区:
生物学4区
文献类型:
--
作者:
Finotello F;Lavezzo E;Bianco L;Barzon L;Mazzon P;Fontana P;Toppo S;Di Camillo B

文献摘要

被引文献

相似文献

在过去的十年里,下一代测序技术被广泛应用于定量转录组学,使RNA测序成为测量和比较基因转录水平的有价值的微阵列的替代方案。虽然已经提出了几种方法来通过数据归一化来提供对转录丰度的无偏估计,但所有这些方法都是基于对每一份转录记录上的总阅读次数的初始计数。如果读数不是沿着序列均匀分布,则该过程在原则上对随机噪声是稳健的,实际上是容易出错的,这确实是由于读映射中的排序错误和歧义而发生的。在这里,我们提出了一种新的方法,称为最大计数,将分配给外显子的表达量化为其每个碱基计数的最大值,并与上面描述的标准方法进行比较,该方法考虑与外显子对齐的总读数。使用多个数据集并考虑几个评估标准来比较这两个指标:独立于基因特定的协变量,如外显子长度和GC含量,在量化真实浓度时的准确性和精密度,以及测量对比对质量变化的稳健性。两种方法都显示出较高的准确度和较低的对GC含量的依赖性。然而,与标准方法相比,Maxcount表达量化较少偏向长外显子。此外,它在低表达时表现出较低的技术可变性,并且对比对质量的变化更健壮。总而言之,我们确认,用标准方法计算的计数取决于总结它们的特征的长度,并且对阅读沿着抄本的不均匀分布很敏感。相反,由于读数的不均匀分布,最大计数对偏差具有很强的鲁棒性,并且具有较低的技术可变性。因此,我们建议将最大计数作为定量RNA测序应用的一种替代方法。
In the last decade, Next-Generation Sequencing technologies have been extensively applied to quantitative transcriptomics, making RNA sequencing a valuable alternative to microarrays for measuring and comparing gene transcription levels. Although several methods have been proposed to provide an unbiased estimate of transcript abundances through data normalization, all of them are based on an initial count of the total number of reads mapping on each transcript. This procedure, in principle robust to random noise, is actually error-prone if reads are not uniformly distributed along sequences, as happens indeed due to sequencing errors and ambiguity in read mapping. Here we propose a new approach, called maxcounts, to quantify the expression assigned to an exon as the maximum of its per-base counts, and we assess its performance in comparison with the standard approach described above, which considers the total number of reads aligned to an exon. The two measures are compared using multiple data sets and considering several evaluation criteria: independence from gene-specific covariates, such as exon length and GC-content, accuracy and precision in the quantification of true concentrations and robustness of measurements to variations of alignments quality. Both measures show high accuracy and low dependency on GC-content. However, maxcounts expression quantification is less biased towards long exons with respect to the standard approach. Moreover, it shows lower technical variability at low expressions and is more robust to variations in the quality of alignments. In summary, we confirm that counts computed with the standard approach depend on the length of the feature they are summarized on, and are sensitive to the non-uniform distribution of reads along transcripts. On the opposite, maxcounts are robust to biases due to the non-uniformity distribution of reads and are characterized by a lower technical variability. Hence, we propose maxcounts as an alternative approach for quantitative RNA-sequencing applications.