A mixture model for expression deconvolution from RNA-seq in heterogeneous tissues.

A mixture model for expression deconvolution from RNA-seq in heterogeneous tissues.
复制标题

DOI:
10.1186/1471-2105-14-s5-s11
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Xie X
Xie X
中科院分区:
生物学4区
文献类型:
--
作者:
Li Y;Xie X

文献摘要

被引文献

相似文献

RNA-SEQ是一种基于下一代测序的转录组分析方法,正迅速成为综合转录丰度估计的首选方法。RNA-SEQ的准确性受样品纯度的影响很大。RNA-SEQ中一个突出的突出问题是如何估计不同组织中的转录丰度,其中一个样本由不止一种细胞类型组成,这种不均质性可能会严重混淆每一种细胞类型的转录丰度估计。尽管已经提出了解剖多种不同细胞类型的实验方法,但计算上的“去卷积”异质组织提供了一个有吸引力的替代方案,因为它保持了组织样本以及随后的分子含量产量的完整性。在这里,我们提出了一种基于概率模型的方法,即混合组织样本的转录估计(TEMT),来从异质组织样本的RNA-SEQ数据中估计每种感兴趣的细胞类型的转录丰度。TEMT结合了位置偏差和特定于序列的偏差,其在线EM算法只需要与数据大小成比例的运行时间和较小的常量内存。我们在模拟数据和最近发布的ENCODE数据上测试了所提出的方法,并表明TEMT显著优于当前未考虑组织异质性的最先进方法。目前,TEMT只能解决由两种细胞类型引起的组织异质性,但可以扩展到处理由多种细胞类型引起的组织异质性。TEMT是用Python语言编写的,可以在https://github.com/uci-cbcl/TEMT.上免费获得本文提出的基于概率模型的方法为从异质组织样本中分析RNA-SEQ数据提供了一种新的方法。通过对模拟数据和ENCODE数据的应用,我们表明,明确考虑组织的异质性可以显著提高转录本丰度估计的准确性。
RNA-seq, a next-generation sequencing based method for transcriptome analysis, is rapidly emerging as the method of choice for comprehensive transcript abundance estimation. The accuracy of RNA-seq can be highly impacted by the purity of samples. A prominent, outstanding problem in RNA-seq is how to estimate transcript abundances in heterogeneous tissues, where a sample is composed of more than one cell type and the inhomogeneity can substantially confound the transcript abundance estimation of each individual cell type. Although experimental methods have been proposed to dissect multiple distinct cell types, computationally "deconvoluting" heterogeneous tissues provides an attractive alternative, since it keeps the tissue sample as well as the subsequent molecular content yield intact. Here we propose a probabilistic model-based approach, Transcript Estimation from Mixed Tissue samples (TEMT), to estimate the transcript abundances of each cell type of interest from RNA-seq data of heterogeneous tissue samples. TEMT incorporates positional and sequence-specific biases, and its online EM algorithm only requires a runtime proportional to the data size and a small constant memory. We test the proposed method on both simulation data and recently released ENCODE data, and show that TEMT significantly outperforms current state-of-the-art methods that do not take tissue heterogeneity into account. Currently, TEMT only resolves the tissue heterogeneity resulting from two cell types, but it can be extended to handle tissue heterogeneity resulting from multi cell types. TEMT is written in python, and is freely available at https://github.com/uci-cbcl/TEMT. The probabilistic model-based approach proposed here provides a new method for analyzing RNA-seq data from heterogeneous tissue samples. By applying the method to both simulation data and ENCODE data, we show that explicitly accounting for tissue heterogeneity can significantly improve the accuracy of transcript abundance estimation.