Statistical Modeling of RNA-Seq Data

Statistical Modeling of RNA-Seq Data
复制标题

DOI:
10.1214/10-sts343
复制
发表时间:
2011-02-01
影响因子:
5.7
通讯作者:
Wong, Wing Hung
Wong, Wing Hung
中科院分区:
数学2区
文献类型:
--
作者:
Salzman, Julia;Jiang, Hui;Wong, Wing Hung

文献摘要

被引文献

相似文献

最近,RNA 超高通量测序 (RNA-Seq) 已被开发为一种分析基因表达的方法。通过获得数千万甚至数亿的转录序列读数,RNA-Seq 实验可以对任何感兴趣样本中的基因(转录本)群体进行全面调查。本文介绍了一种统计模型,用于根据 RNA-Seq 数据估计同种型丰度,并且足够灵活,可以适应单端和配对端 RNA-Seq 数据以及沿转录本长度的采样偏差。基于模型的最小充分统计量的推导,提供了模型的最大似然估计器的计算上可行的实现。此外,结果表明,在固定测序深度下,使用配对末端 RNA-Seq 比单末端测序可提供更准确的异构体丰度估计。还给出了模拟研究。
Recently, ultra high-throughput sequencing of RNA (RNA-Seq) has been developed as an approach for analysis of gene expression. By obtaining tens or even hundreds of millions of reads of transcribed sequences, an RNA-Seq experiment can offer a comprehensive survey of the population of genes (transcripts) in any sample of interest. This paper introduces a statistical model for estimating isoform abundance from RNA-Seq data and is flexible enough to accommodate both single end and paired end RNA-Seq data and sampling bias along the length of the transcript. Based on the derivation of minimal sufficient statistics for the model, a computationally feasible implementation of the maximum likelihood estimator of the model is provided. Further, it is shown that using paired end RNA-Seq provides more accurate isoform abundance estimates than single end sequencing at fixed sequencing depth. Simulation studies are also given.