Challenges and strategies in transcriptome assembly and differential gene expression quantification. A comprehensive in silico assessment of RNA-seq experiments

Challenges and strategies in transcriptome assembly and differential gene expression quantification. A comprehensive in silico assessment of RNA-seq experiments
复制标题

DOI:
10.1111/mec.12014
复制
发表时间:
2013-02-01
期刊:
影响因子:
4.9
通讯作者:
Wolf, Jochen B. W.
Wolf, Jochen B. W.
中科院分区:
生物学1区
文献类型:
--
作者:
Vijay, Nagarjun;Poelstra, Jelmer W.;Wolf, Jochen B. W.

文献摘要

被引文献

相似文献

转录组鸟枪法测序 (RNA-seq) 很容易受到遗传学家和分子生态学家的欢迎。与所有高通量技术一样,了解哪些分析策略最适合以及哪些参数可能会影响数据解释至关重要。在这里,我们使用全面的模拟方法来探索转录组的各种特征(复杂性、多态性程度 p、选择性剪接)、技术处理(测序误差 e、文库标准化)和生物信息工作流程(从头到图组装、参考基因组质量)如何影响转录组质量和差异基因表达 (DE) 的推断。我们发现转录组组装和基因表达谱分析(EdgeR 与 BaySeq 软件)即使在没有参考基因组的情况下也能很好地工作,并且在广泛的参数范围内都具有稳健性。我们建议不要进行文库标准化,并且在大多数情况下主张将组件映射到不同姐妹进化枝的注释基因组,这通常优于从头组装(Trans-Abyss、Trinity、Soapdenovo-Trans)。转录组复杂性(大小、旁系同源物、选择性剪接亚型)对组装和 DE 分析产生负面影响,而测序错误和多态性的影响几乎可以忽略不计。最后,我们强调了从头组装的基因名称分配的挑战、作图策略的重要性,并提高了对与参考基因组质量相关的挑战的认识。总的来说,我们的结果具有重要的实践和方法学意义,可以为 RNA-seq 实验的设计和分析提供指导,特别是对于缺乏基因组背景信息的生物体。
Transcriptome Shotgun Sequencing (RNA-seq) has been readily embraced by geneticists and molecular ecologists alike. As with all high-throughput technologies, it is critical to understand which analytic strategies are best suited and which parameters may bias the interpretation of the data. Here we use a comprehensive simulation approach to explore how various features of the transcriptome (complexity, degree of polymorphism p, alternative splicing), technological processing (sequencing error e, library normalization) and bioinformatic workflow (de novo vs. mapping assembly, reference genome quality) impact transcriptome quality and inference of differential gene expression (DE). We find that transcriptome assembly and gene expression profiling (EdgeR vs. BaySeq software) works well even in the absence of a reference genome and is robust across a broad range of parameters. We advise against library normalization and in most situations advocate mapping assemblies to an annotated genome of a divergent sister clade, which generally outperformed de novo assembly (Trans-Abyss, Trinity, Soapdenovo-Trans). Transcriptome complexity (size, paralogs, alternative splicing isoforms) negatively affected the assembly and DE profiling, whereas the effects of sequencing error and polymorphism were almost negligible. Finally, we highlight the challenge of gene name assignment for de novo assemblies, the importance of mapping strategies and raise awareness of challenges associated with the quality of reference genomes. Overall, our results have significant practical and methodological implications and can provide guidance in the design and analysis of RNA-seq experiments, particularly for organisms where genomic background information is lacking.