Sources of variation in cell-type RNA-Seq profiles.

Sources of variation in cell-type RNA-Seq profiles.
复制标题

DOI:
10.1371/journal.pone.0239495
复制
发表时间:
2020
期刊:
影响因子:
3.7
通讯作者:
Nielsen J
Nielsen J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Gustafsson J;Held F;Robinson JL;Björnson E;Jörnsten R;Nielsen J

文献摘要

参考文献

被引文献

相似文献

在大量RNA-Seq样品上操作的许多计算方法都需要细胞类型特异性基因表达谱,例如细胞类型部分的反褶积和数字细胞术。然而,由于技术因素和细胞状态和环境的生物学差异,一种细胞类型的基因表达谱可能会发生很大的变化,从而降低了这种方法的有效性。在这里,我们调查了哪些因素对这种变化贡献最大。我们评估了不同的归一化方法,量化了不同因素解释的方差,评估了细胞类型分数对反褶积的影响,并检查了基于uni的单细胞RNA-Seq和大量RNA-Seq之间的差异。我们调查了包含B细胞和T细胞的公开可获得的大量和单细胞RNA-Seq数据集,发现实验室之间的技术差异是实质性的,甚至对于专门选择用于反褶积的基因,这种差异对反褶积具有混淆效应。来源组织也是一个重要因素,这突出了使用来自血液和其他组织混合物的细胞类型谱的挑战。我们还表明,基于uni的单细胞和大量RNA-Seq方法之间的许多差异可以通过单细胞样品中每个mRNA分子的读取重复数来解释。我们的工作表明,在创建将与大量样品一起使用的细胞类型特定基因表达谱时,匹配或纠正技术因素的重要性。
Cell-type specific gene expression profiles are needed for many computational methods operating on bulk RNA-Seq samples, such as deconvolution of cell-type fractions and digital cytometry. However, the gene expression profile of a cell type can vary substantially due to both technical factors and biological differences in cell state and surroundings, reducing the efficacy of such methods. Here, we investigated which factors contribute most to this variation. We evaluated different normalization methods, quantified the variance explained by different factors, evaluated the effect on deconvolution of cell type fractions, and examined the differences between UMI-based single-cell RNA-Seq and bulk RNA-Seq. We investigated a collection of publicly available bulk and single-cell RNA-Seq datasets containing B and T cells, and found that the technical variation across laboratories is substantial, even for genes specifically selected for deconvolution, and this variation has a confounding effect on deconvolution. Tissue of origin is also a substantial factor, highlighting the challenge of using cell type profiles derived from blood with mixtures from other tissues. We also show that much of the differences between UMI-based single-cell and bulk RNA-Seq methods can be explained by the number of read duplicates per mRNA molecule in the single-cell sample. Our work shows the importance of either matching or correcting for technical factors when creating cell-type specific gene expression profiles that are to be used together with bulk samples.
DOI: 10.1186/s13059-016-0947-7
发表时间: 2016-04-27
期刊: Genome biology
影响因子: 12.3
作者:
Lun AT;Bach K;Marioni JC
通讯作者: Marioni JC
DOI: 10.1186/s12859-016-1323-z
发表时间: 2016-11-25
期刊: BMC bioinformatics
影响因子: 3
作者:
Hoffman GE;Schadt EE
通讯作者: Schadt EE
DOI: 10.1038/s41587-020-0465-8
发表时间: 2020-04-06
影响因子: 46.9
作者:
Ding, Jiarui;Adiconis, Xian;Levin, Joshua Z.
通讯作者: Levin, Joshua Z.
DOI: 10.1093/biostatistics/kxx028
发表时间: 2018-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Hicks, Stephanie C.;Okrah, Kwame;Bravo, Hector Corrada
通讯作者: Bravo, Hector Corrada
DOI: 10.1371/journal.pone.0006098
发表时间: 2009-07-01
期刊: PloS one
影响因子: 3.7
作者:
Abbas AR;Wolslegel K;Seshasayee D;Modrusan Z;Clark HF
通讯作者: Clark HF