SQANTI: extensive characterization of long-read transcript sequences for quality control in full-length transcriptome identification and quantification.

SQANTI: extensive characterization of long-read transcript sequences for quality control in full-length transcriptome identification and quantification.
复制标题

DOI:
10.1101/gr.222976.117
复制
发表时间:
2018-03-01
期刊:
影响因子:
7
通讯作者:
Conesa A
Conesa A
中科院分区:
生物学1区
文献类型:
--
作者:
Tardaguila M;de la Fuente L;Marti C;Pereira C;Pardo-Palacios FJ;Del Risco H;Ferrell M;Mellado M;Macchietto M;Verheggen K;Edelmann M;Ezkurdia I;Vazquez J;Tress M;Mortazavi A;Martens L;Rodriguez-Navarro S;Moreno-Manzano V;Conesa A

文献摘要

参考文献

被引文献

相似文献

使用长阅读对全长转录本进行高通量测序,为发现数千个新的转录本铺平了道路,即使是在注释良好的哺乳动物物种中也是如此。测序技术的进步产生了对能够表征这些新变种的研究和工具的需求。在这里,我们介绍了SQANTI,一个用于长读记录分类的自动化流水线,可以评估数据的质量,以及使用47个唯一描述符的预处理流水线。我们将SQANTI应用于使用太平洋生物科学(PacBio)的神经元小鼠转录组(PacBio),并说明了该工具如何有效地表征和描述全长转录组的组成。我们通过PCR对豆腐PacBio转录本进行了广泛的评估,发现许多新的转录本是测序方法的技术产物,并且SQANTI质量描述符可以用来设计过滤策略来删除它们。在这个精选的转录组中,大多数新的转录本是现有剪接位点的新组合,导致新的ORF比新的UTRs更频繁,并具有丰富的一般代谢和神经特异性功能。我们表明,这些新的转录本对基于最先进的短读量化算法的转录本水平的正确量化具有重大影响。通过将我们的等转录组与公共蛋白质组学数据库进行比较,我们发现替代的异构体对于蛋白质组学检测是难以捉摸的。SQANTI允许用户通过提供工具来提供经过质量评估和精选的全长抄本,从而最大限度地提高长篇阅读技术的分析结果。
High-throughput sequencing of full-length transcripts using long reads has paved the way for the discovery of thousands of novel transcripts, even in well-annotated mammalian species. The advances in sequencing technology have created a need for studies and tools that can characterize these novel variants. Here, we present SQANTI, an automated pipeline for the classification of long-read transcripts that can assess the quality of data and the preprocessing pipeline using 47 unique descriptors. We apply SQANTI to a neuronal mouse transcriptome using Pacific Biosciences (PacBio) long reads and illustrate how the tool is effective in characterizing and describing the composition of the full-length transcriptome. We perform extensive evaluation of ToFU PacBio transcripts by PCR to reveal that an important number of the novel transcripts are technical artifacts of the sequencing approach and that SQANTI quality descriptors can be used to engineer a filtering strategy to remove them. Most novel transcripts in this curated transcriptome are novel combinations of existing splice sites, resulting more frequently in novel ORFs than novel UTRs, and are enriched in both general metabolic and neural-specific functions. We show that these new transcripts have a major impact in the correct quantification of transcript levels by state-of-the-art short-read-based quantification algorithms. By comparing our iso-transcriptome with public proteomics databases, we find that alternative isoforms are elusive to proteogenomics detection. SQANTI allows the user to maximize the analytical outcome of long-read technologies by providing the tools to deliver quality-evaluated and curated full-length transcriptomes.
DOI: 10.1093/database/bas014
发表时间: 2012
期刊: Database : the journal of biological databases and curation
影响因子: --
作者:
Frankish A;Mudge JM;Thomas M;Harrow J
通讯作者: Harrow J
DOI: 10.1002/glia.22900
发表时间: 2016-01
期刊: Glia
影响因子: 6.2
作者:
Amaral AI;Hadera MG;Tavares JM;Kotter MR;Sonnewald U
通讯作者: Sonnewald U
DOI: 10.1371/journal.pone.0157779
发表时间: 2016
期刊: PloS one
影响因子: 3.7
作者:
Cartolano M;Huettel B;Hartwig B;Reinhardt R;Schneeberger K
通讯作者: Schneeberger K
DOI: 10.1038/ncomms11706
发表时间: 2016-06-24
影响因子: 16.6
作者:
Abdel-Ghany SE;Hamilton M;Jacobi JL;Ngam P;Devitt N;Schilkey F;Ben-Hur A;Reddy AS
通讯作者: Reddy AS
人蛋白质组项目质谱数据解释指南2.1。
DOI: 10.1021/acs.jproteome.6b00392
发表时间: 2016-11-04
影响因子: 4.4
作者:
Deutsch EW;Overall CM;Van Eyk JE;Baker MS;Paik YK;Weintraub ST;Lane L;Martens L;Vandenbrouck Y;Kusebauch U;Hancock WS;Hermjakob H;Aebersold R;Moritz RL;Omenn GS
通讯作者: Omenn GS