Variability in donor leukocyte counts confound the use of common RNA sequencing data normalization strategies in transcriptomic biomarker studies performed with whole blood.

Variability in donor leukocyte counts confound the use of common RNA sequencing data normalization strategies in transcriptomic biomarker studies performed with whole blood.
复制标题

DOI:
10.1038/s41598-023-41443-4
复制
发表时间:
2023-09-19
期刊:
影响因子:
4.6
通讯作者:
--
中科院分区:
综合性期刊3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

通过下一代测序从全血产生的基因表达数据经常用于旨在鉴定基于mRNA的生物标志物组的研究,其可用于诊断或监测人类疾病。这些研究通常采用数据标准化技术,更通常用于分析来自实体组织的数据,这在很大程度上是在标本具有相似转录组组成的一般假设下进行的。然而,当使用从全血生成的数据时,可能会违反这一假设,全血更具细胞动态性,从而导致潜在的混淆。在这项研究中,我们使用下一代测序结合流式细胞术来评估供体白细胞计数对从138名人类受试者的队列中采样的全血标本的转录组成的影响,然后随后检查了四种常用的数据归一化方法对我们检测标本间生物学差异的能力的影响,使用流式细胞术数据对每个样本的真实细胞和分子身份进行基准测试。来自具有不同白细胞计数的供体的全血样品在转录物丰度的全基因组分布和基因水平表达模式方面表现出显著差异。因此,我们测试的三种标准化策略,包括中位数比(MRN),m值的修剪平均值(TMM)和分位数标准化,明显掩盖了数据的真实生物学结构,削弱了我们检测mRNA水平真实样本间差异的能力。提高我们检测真实生物学差异的能力的唯一策略是通过测序深度简单缩放读段计数,与上述方法不同,它不对转录组组成进行假设。
Gene expression data generated from whole blood via next generation sequencing is frequently used in studies aimed at identifying mRNA-based biomarker panels with utility for diagnosis or monitoring of human disease. These investigations often employ data normalization techniques more typically used for analysis of data originating from solid tissues, which largely operate under the general assumption that specimens have similar transcriptome composition. However, this assumption may be violated when working with data generated from whole blood, which is more cellularly dynamic, leading to potential confounds. In this study, we used next generation sequencing in combination with flow cytometry to assess the influence of donor leukocyte counts on the transcriptional composition of whole blood specimens sampled from a cohort of 138 human subjects, and then subsequently examined the effect of four frequently used data normalization approaches on our ability to detect inter-specimen biological variance, using the flow cytometry data to benchmark each specimens true cellular and molecular identity. Whole blood samples originating from donors with differing leukocyte counts exhibited dramatic differences in both genome-wide distributions of transcript abundance and gene-level expression patterns. Consequently, three of the normalization strategies we tested, including median ratio (MRN), trimmed mean of m-values (TMM), and quantile normalization, noticeably masked the true biological structure of the data and impaired our ability to detect true interspecimen differences in mRNA levels. The only strategy that improved our ability to detect true biological variance was simple scaling of read counts by sequencing depth, which unlike the aforementioned approaches, makes no assumptions regarding transcriptome composition.
DOI: 10.1038/npjgenmed.2016.38
发表时间: 2016
影响因子: 5.3
作者:
O'Connell GC;Petrone AB;Treadway MB;Tennant CS;Lucke-Wold N;Chantler PD;Barr TL
通讯作者: Barr TL
WGCNA:用于加权相关网络分析的 R 包。
DOI: 10.1186/1471-2105-9-559
发表时间: 2008-12-29
期刊: BMC bioinformatics
影响因子: 3
作者:
Langfelder P;Horvath S
通讯作者: Horvath S
DOI: 10.1038/nature09247
发表时间: 2010-08-19
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1038/s41598-020-59516-z
发表时间: 2020-02-17
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Arora, Sonali;Pattwell, Siobhan S.;Bolouri, Hamid
通讯作者: Bolouri, Hamid
标准化如何影响 RNA-seq 疾病诊断?
DOI: 10.1016/j.jbi.2018.07.016
发表时间: 2018-09-01
影响因子: 4.5
作者:
Han, Henry;Men, Ke
通讯作者: Men, Ke