Estimating cell type composition using isoform expression one gene at a time.

Estimating cell type composition using isoform expression one gene at a time.
复制标题

一次使用一个基因的异构体表达来估计细胞类型组成。

DOI:
10.1111/biom.13614
复制
发表时间:
2023
期刊:
影响因子:
1.9
通讯作者:
Ibrahim,JosephG
Ibrahim,JosephG
中科院分区:
数学3区
文献类型:
--
作者:
Heiling,HillaryM;Wilson,DouglasR;Rashid,NaimU;Sun,Wei;Ibrahim,JosephG

文献摘要

相似文献

人体组织样本通常是异质细胞类型的混合物,这可能会混淆来自这些组织的基因表达数据的分析。组织样本的细胞类型组成本身可能是感兴趣的,并且需要对差异基因表达进行适当的分析。各种计算方法已经开发估计细胞类型比例使用基因水平的表达数据。然而,RNA同型异构体也可以在不同细胞类型中表达差异,同型异构体水平的表达可以与基因水平的表达一样或更多地用于确定细胞类型起源。我们提出了一种新的计算方法,IsoDeconvMM,它使用同种异构体水平的基因表达数据来估计细胞类型分数。IsoDeconvMM的一个新颖而有用的功能是,它可以仅使用单个基因来估计细胞类型比例,尽管在实践中,我们建议将几十个基因的估计汇总起来以获得更准确的结果。我们使用一个独特的数据集来展示IsoDeconvMM的性能,该数据集包含超过135个个体的细胞类型特异性RNA-seq数据。该数据集允许我们评估不同的方法给定的生物变异细胞类型特异性基因表达数据在个体之间。我们用额外的模拟进一步补充了这一分析。
Human tissue samples are often mixtures of heterogeneous cell types, which can confound the analyses of gene expression data derived from such tissues. The cell type composition of a tissue sample may itself be of interest and is needed for proper analysis of differential gene expression. A variety of computational methods have been developed to estimate cell type proportions using gene-level expression data. However, RNA isoforms can also be differentially expressed across cell types, and isoform-level expression could be equally or more informative for determining cell type origin than gene-level expression. We propose a new computational method, IsoDeconvMM, which estimates cell type fractions using isoform-level gene expression data. A novel and useful feature of IsoDeconvMM is that it can estimate cell type proportions using only a single gene, though in practice we recommend aggregating estimates of a few dozen genes to obtain more accurate results. We demonstrate the performance of IsoDeconvMM using a unique data set with cell type–specific RNA-seq data across more than 135 individuals. This data set allows us to evaluate different methods given the biological variation of cell type–specific gene expression data across individuals. We further complement this analysis with additional simulations.