Smooth quantile normalization

Smooth quantile normalization
复制标题

DOI:
10.1093/biostatistics/kxx028
复制
发表时间:
2018-04-01
期刊:
影响因子:
2.1
通讯作者:
Bravo, Hector Corrada
Bravo, Hector Corrada
中科院分区:
数学2区
文献类型:
--
作者:
Hicks, Stephanie C.;Okrah, Kwame;Bravo, Hector Corrada

文献摘要

被引文献

相似文献

样本间标准化是基因组数据分析中的关键步骤,以消除高通量数据中的系统偏差和不必要的技术变化。全局归一化方法基于这样的假设,即观察到的全局特性的变异性是由于技术原因,与感兴趣的生物学无关。例如,一些方法通过缩放特征以在样品之间具有相似的中值来校正测序读段计数的差异,但这些方法未能减少其他形式的不需要的技术变化。分位数归一化等方法将样本之间的统计分布转换为相同的,并假设分布中的全局差异仅由技术变化引起。然而,如果这些假设被违反,例如,如果生物条件或组之间的统计分布存在全局差异,并且外部信息(如阴性或对照特征)不可用,则如何进行标准化仍不清楚。在这里,我们介绍了分位数归一化的一般化,称为平滑分位数归一化(qsmooth),它基于以下假设:每个样本的统计分布在生物组或条件内应该是相同的(或具有相同的分布形状),但允许它们在组之间可能不同。我们说明了我们的方法在几个高通量数据集上的优势,这些数据集具有对应于不同生物条件的分布的全局差异。我们还进行了Monte Carlo模拟研究,以说明与其他全局归一化方法相比,qsmooth的偏差-方差权衡和均方根误差。可从https://github.com/stephaniehicks/gsmooth获得软件实现。
Between-sample normalization is a critical step in genomic data analysis to remove systematic bias and unwanted technical variation in high-throughput data. Global normalization methods are based on the assumption that observed variability in global properties is due to technical reasons and are unrelated to the biology of interest. For example, some methods correct for differences in sequencing read counts by scaling features to have similar median values across samples, but these fail to reduce other forms of unwanted technical variation. Methods such as quantile normalization transform the statistical distributions across samples to be the same and assume global differences in the distribution are induced by only technical variation. However, it remains unclear how to proceed with normalization if these assumptions are violated, for example, if there are global differences in the statistical distributions between biological conditions or groups, and external information, such as negative or control features, is not available. Here, we introduce a generalization of quantile normalization, referred to as smooth quantile normalization (qsmooth), which is based on the assumption that the statistical distribution of each sample should be the same (or have the same distributional shape) within biological groups or conditions, but allowing that they may differ between groups. We illustrate the advantages of our method on several high-throughput datasets with global differences in distributions corresponding to different biological conditions. We also perform a Monte Carlo simulation study to illustrate the bias-variance tradeoff and root mean squared error of qsmooth compared to other global normalization methods. A software implementation is available from https://github.com/stephaniehicks/gsmooth.