ComBat-seq: batch effect adjustment for RNA-seq count data.

ComBat-seq: batch effect adjustment for RNA-seq count data.
复制标题

DOI:
10.1093/nargab/lqaa078
复制
发表时间:
2020-09
影响因子:
4.6
通讯作者:
Johnson WE
Johnson WE
中科院分区:
其他
文献类型:
--
作者:
Zhang Y;Parmigiani G;Johnson WE

文献摘要

被引文献

相似文献

批处理效应或由于跨批处理的技术因素差异而导致的数据差异通常会阻碍基因组数据批次增加统计能力的批处理的好处。因此,至关重要的是,有效解决基因组数据中的批处理效应以克服这些挑战。许多现有的批处理效果调整方法都假定数据遵循连续的,钟形的高斯分布。但是,在RNA-seq研究中,数据通常偏斜,过度分散的计数,因此该假设不合适,可能会导致错误的结果。负二项式回归模型以前已被用于更好地捕获计数的特性。我们使用负二项式回归模型开发了一种批化方法,即战斗式,该模型保留了RNA-seq研究中计数数据的整数性质,从而使批次调整后的数据与需要整数数量计数的常见微分表达软件包兼容。我们在现实的模拟中表明,与其他可用方法调整的数据相比,战斗seq调整后的数据可以更好地统计能力和差异表达中的假阳性的控制。我们在一个真实的数据示例中进一步证明了战斗seq成功地消除了批处理效应并恢复数据中的生物学信号。
The benefit of integrating batches of genomic data to increase statistical power is often hindered by batch effects, or unwanted variation in data caused by differences in technical factors across batches. It is therefore critical to effectively address batch effects in genomic data to overcome these challenges. Many existing methods for batch effects adjustment assume the data follow a continuous, bell-shaped Gaussian distribution. However in RNA-seq studies the data are typically skewed, over-dispersed counts, so this assumption is not appropriate and may lead to erroneous results. Negative binomial regression models have been used previously to better capture the properties of counts. We developed a batch correction method, ComBat-seq, using a negative binomial regression model that retains the integer nature of count data in RNA-seq studies, making the batch adjusted data compatible with common differential expression software packages that require integer counts. We show in realistic simulations that the ComBat-seq adjusted data results in better statistical power and control of false positives in differential expression compared to data adjusted by the other available methods. We further demonstrated in a real data example that ComBat-seq successfully removes batch effects and recovers the biological signal in the data.