TIGAR2: sensitive and accurate estimation of transcript isoform expression with longer RNA-Seq reads.

TIGAR2: sensitive and accurate estimation of transcript isoform expression with longer RNA-Seq reads.
复制标题

DOI:
10.1186/1471-2164-15-s10-s5
复制
发表时间:
2014
期刊:
影响因子:
4.4
通讯作者:
Nagasaki M
Nagasaki M
中科院分区:
生物学2区
文献类型:
--
作者:
Nariai N;Kojima K;Mimori T;Sato Y;Kawai Y;Yamaguchi-Kabata Y;Nagasaki M

文献摘要

被引文献

相似文献

高通量RNA测序(RNA- seq)能够在单碱基分辨率下定量和鉴定转录本。最近,由于新型测序技术的发展以及下一代测序仪化学试剂的改进,更长的序列读数变得可用。虽然已经提出了几种计算方法来从RNA-Seq数据中量化基因表达水平,但它们对于较长的读数(例如> 250 bp)还没有充分优化。我们提出了TIGAR2,这是一种从固定长度和可变长度RNA-Seq数据中定量转录物异构体的统计方法。我们的方法基于reads与参考cDNA序列的间隙比对来模拟序列的替换、缺失和插入错误,以便将敏感的read-aligners(如Bowtie2和BWA-MEM)有效地纳入我们的产品线。同时,在变分贝叶斯推理中引入了启发式算法,提高了计算速度。我们将TIGAR2应用于人类样本的模拟数据和真实数据,并与现有方法进行比较,评估使用TIGAR2进行转录本定量的性能。TIGAR2是从RNA-Seq数据中定量转录异构体丰度的灵敏而准确的工具。该方法在固定长度(单端和对端分别为100 bp、250 bp、500 bp和1000 bp)和可变长度(特别是长度大于250 bp)上优于现有方法。
High-throughput RNA sequencing (RNA-Seq) enables quantification and identification of transcripts at single-base resolution. Recently, longer sequence reads become available thanks to the development of new types of sequencing technologies as well as improvements in chemical reagents for the Next Generation Sequencers. Although several computational methods have been proposed for quantifying gene expression levels from RNA-Seq data, they are not sufficiently optimized for longer reads (e.g. > 250 bp). We propose TIGAR2, a statistical method for quantifying transcript isoforms from fixed and variable length RNA-Seq data. Our method models substitution, deletion, and insertion errors of sequencers based on gapped-alignments of reads to the reference cDNA sequences so that sensitive read-aligners such as Bowtie2 and BWA-MEM are effectively incorporated in our pipeline. Also, a heuristic algorithm is implemented in variational Bayesian inference for faster computation. We apply TIGAR2 to both simulation data and real data of human samples and evaluate performance of transcript quantification with TIGAR2 in comparison to existing methods. TIGAR2 is a sensitive and accurate tool for quantifying transcript isoform abundances from RNA-Seq data. Our method performs better than existing methods for the fixed-length reads (100 bp, 250 bp, 500 bp, and 1000 bp of both single-end and paired-end) and variable-length reads, especially for reads longer than 250 bp.