Overcoming the impacts of two-step batch effect correction on gene expression estimation and inference.

Overcoming the impacts of two-step batch effect correction on gene expression estimation and inference.
复制标题

DOI:
10.1093/biostatistics/kxab039
复制
发表时间:
2023-07-14
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

Nonignorable technical variation is commonly observed across data from multiple experimental runs, platforms, or studies. These so-called batch effects can lead to difficulty in merging data from multiple sources, as they can severely bias the outcome of the analysis. Many groups have developed approaches for removing batch effects from data, usually by accommodating batch variables into the analysis (one-step correction) or by preprocessing the data prior to the formal or final analysis (two-step correction). One-step correction is often desirable due it its simplicity, but its flexibility is limited and it can be difficult to include batch variables uniformly when an analysis has multiple stages. Two-step correction allows for richer models of batch mean and variance. However, prior investigation has indicated that two-step correction can lead to incorrect statistical inference in downstream analysis. Generally speaking, two-step approaches introduce a correlation structure in the corrected data, which, if ignored, may lead to either exaggerated or diminished significance in downstream applications such as differential expression analysis. Here, we provide more intuitive and more formal evaluations of the impacts of two-step batch correction compared to existing literature. We demonstrate that the undesired impacts of two-step correction (exaggerated or diminished significance) depend on both the nature of the study design and the batch effects. We also provide strategies for overcoming these negative impacts in downstream analyses using the estimated correlation matrix of the corrected data. We compare the results of our proposed workflow with the results from other published one-step and two-step methods and show that our methods lead to more consistent false discovery controls and power of detection across a variety of batch effect scenarios. Software for our method is available through GitHub (https://github.com/jtleek/sva-devel) and will be available in future versions of the R package in the Bioconductor project (https://bioconductor.org/packages/release/bioc/html/sva.html).
DOI: 10.1080/14029251.2013.855050
发表时间: 2013-09-01
影响因子: 0.7
作者:
Zusmanovich, Pasha
通讯作者: Zusmanovich, Pasha
DOI: 10.1093/bioinformatics/btw538
发表时间: 2016-12-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Manimaran S;Selby HM;Okrah K;Ruberman C;Leek JT;Quackenbush J;Haibe-Kains B;Bravo HC;Johnson WE
通讯作者: Johnson WE
通过替代变量分析捕获基因表达研究中的异质性。
DOI: 10.1371/journal.pgen.0030161
发表时间: 2007-09
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Leek, Jeffrey T.;Storey, John D.
通讯作者: Storey, John D.
DOI: 10.1016/j.tube.2018.01.002
发表时间: 2018-03-01
期刊: TUBERCULOSIS
影响因子: 3.2
作者:
Leong, Samantha;Zhao, Yue;Salgame, Padmini
通讯作者: Salgame, Padmini
DOI: 10.1093/biostatistics/kxr034
发表时间: 2012-07-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Gagnon-Bartsch, Johann A.;Speed, Terence P.
通讯作者: Speed, Terence P.