Sparse linear modeling of next-generation mRNA sequencing (RNA-Seq) data for isoform discovery and abundance estimation

Sparse linear modeling of next-generation mRNA sequencing (RNA-Seq) data for isoform discovery and abundance estimation
复制标题

DOI:
10.1073/pnas.1113972108
复制
发表时间:
2011-12-13
影响因子:
11.1
通讯作者:
Bickel, Peter J.
Bickel, Peter J.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Li, Jingyi Jessica;Jiang, Ci-Ren;Bickel, Peter J.

文献摘要

被引文献

相似文献

自从下一代mRNA测序(RNA-Seq)技术开始以来,已经进行了各种尝试以利用RNA-Seq数据从头组装全长mRNA同种型并估计同种型的丰度。然而,对于具有多于几个外显子的基因,问题往往是具有挑战性的,并且通常涉及统计建模中的可识别性问题。我们开发了一种称为“用于异构体发现和丰度估计的RNA-Seq数据的稀疏线性建模”(SLIDE)的统计方法,该方法将外显子边界和RNA-Seq数据作为输入,以辨别最有可能存在于RNA-Seq样品中的mRNA异构体集。SLIDE基于线性模型,其设计矩阵对来自不同mRNA同种型的RNA-Seq读数的采样概率进行建模。为了解决模型的不可识别性问题,SLIDE使用了一个修改的Lasso过程进行参数估计。与确定性同种型组装算法(例如,Cufflinks),SLIDE考虑了来自不同亚型的外显子中RNA-Seq读数的随机方面,从而增加了检测更多新亚型的能力。SLIDE的另一个优点是它可以灵活地将其他转录组数据(如RACE,CAGE和EST)纳入其模型,以进一步提高异构体发现的准确性。SLIDE还可以在其他RNA-Seq组装算法的下游工作,以整合新发现的基因和外显子。除了异构体发现之外,SLIDE还依次使用相同的线性模型来估计已发现异构体的丰度。模拟和真实的数据研究表明,SLIDE在异构体发现和丰度估计方面的表现与主要竞争对手一样好或更好。SLIDE软件包可在https://sites.google.com/site/jingyijli/SLIDE.zip上获得。
Since the inception of next-generation mRNA sequencing (RNA-Seq) technology, various attempts have been made to utilize RNA-Seq data in assembling full-length mRNA isoforms de novo and estimating abundance of isoforms. However, for genes with more than a few exons, the problem tends to be challenging and often involves identifiability issues in statistical modeling. We have developed a statistical method called "sparse linear modeling of RNA-Seq data for isoform discovery and abundance estimation" (SLIDE) that takes exon boundaries and RNA-Seq data as input to discern the set of mRNA isoforms that are most likely to present in an RNA-Seq sample. SLIDE is based on a linear model with a design matrix that models the sampling probability of RNA-Seq reads from different mRNA isoforms. To tackle the model unidentifiability issue, SLIDE uses a modified Lasso procedure for parameter estimation. Compared with deterministic isoform assembly algorithms (e.g., Cufflinks), SLIDE considers the stochastic aspects of RNA-Seq reads in exons from different isoforms and thus has increased power in detecting more novel isoforms. Another advantage of SLIDE is its flexibility of incorporating other transcriptomic data such as RACE, CAGE, and EST into its model to further increase isoform discovery accuracy. SLIDE can also work downstream of other RNA-Seq assembly algorithms to integrate newly discovered genes and exons. Besides isoform discovery, SLIDE sequentially uses the same linear model to estimate the abundance of discovered isoforms. Simulation and real data studies show that SLIDE performs as well as or better than major competitors in both isoform discovery and abundance estimation. The SLIDE software package is available at https://sites.google.com/site/jingyijli/SLIDE.zip.