IntAPT: integrated assembly of phenotype-specific transcripts from multiple RNA-seq profiles.

IntAPT: integrated assembly of phenotype-specific transcripts from multiple RNA-seq profiles.
复制标题

IntAPT:来自多个 RNA-seq 配置文件的表型特异性转录本的集成组装。

DOI:
10.1093/bioinformatics/btaa852
复制
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Xuan,Jianhua
Xuan,Jianhua
中科院分区:
--
文献类型:
--
作者:
Shi,Xu;Neuwald,AndrewF;Wang,Xiao;Wang,Tian-Li;Hilakivi-Clarke,Leena;Clarke,Robert;Xuan,Jianhua

文献摘要

相似文献

高通量RNA测序彻底改变了转录组分析的范围和深度。由于RNA-seq数据的噪声和可变性,表型特异性转录组的准确重建具有挑战性。这需要计算识别的成绩单从多个样品相同的表型,鉴于基本的共识转录structure.ResultsWe提出了一种贝叶斯方法,集成组装的表型特异性转录本(IntAPT),识别表型特异性异构体从多个RNA-seq配置文件。IntAPT具有一种新型的双层贝叶斯模型,用于捕获组层中异构体的存在,并量化样品层中异构体的丰度。一个尖峰和平板先验被用来模拟异构体表达,并加强表达的异构体的稀疏性。异构体的存在和它们的表达之间的相似性明确建模,以便于参数估计。使用Gibbs采样迭代估计模型参数以推断联合后验分布,由此可以可靠地确定异构体的存在和丰度。使用模拟和真实的数据集的研究表明,IntAPT始终优于现有的IntAPT方法。实验结果表明,尽管测序错误,IntAPT在多个样品中表现出稳健的性能,从而显著提高了对低丰度表达同种型的鉴定。http://github.com/henryxushi/IntAPT.Supplementary
MotivationHigh-throughput RNA sequencing has revolutionized the scope and depth of transcriptome analysis. Accurate reconstruction of a phenotype-specific transcriptome is challenging due to the noise and variability of RNA-seq data. This requires computational identification of transcripts from multiple samples of the same phenotype, given the underlying consensus transcript structure.ResultsWe present a Bayesian method, integrated assembly of phenotype-specific transcripts (IntAPT), that identifies phenotype-specific isoforms from multiple RNA-seq profiles. IntAPT features a novel two-layer Bayesian model to capture the presence of isoforms at the group layer and to quantify the abundance of isoforms at the sample layer. A spike-and-slab prior is used to model the isoform expression and to enforce the sparsity of expressed isoforms. Dependencies between the existence of isoforms and their expression are modeled explicitly to facilitate parameter estimation. Model parameters are estimated iteratively using Gibbs sampling to infer the joint posterior distribution, from which the presence and abundance of isoforms can reliably be determined. Studies using both simulations and real datasets show that IntAPT consistently outperforms existing methods for the IntAPT. Experimental results demonstrate that, despite sequencing errors, IntAPT exhibits a robust performance among multiple samples, resulting in notably improved identification of expressed isoforms of low abundance.Availability and implementationThe IntAPT package is available at http://github.com/henryxushi/IntAPT.Supplementary informationSupplementary data are available atBioinformaticsonline.