De novo transcriptome assembly of RNA-Seq reads with different strategies

De novo transcriptome assembly of RNA-Seq reads with different strategies
复制标题

DOI:
10.1007/s11427-011-4256-9
复制
发表时间:
2011-12-01
影响因子:
9.1
通讯作者:
Shi TieLiu
Shi TieLiu
中科院分区:
生物学1区
文献类型:
--
作者:
Chen Geng;Yin KangPing;Shi TieLiu

文献摘要

被引文献

相似文献

重新组装转录组是RNA-Seq数据分析中的一种重要方法,它可以帮助我们在没有参考基因组序列的情况下重建转录组和研究基因表达谱。我们分别用人脑和细胞系产生的两个RNA-Seq数据集进行了转录组组装。然后,我们使用三种不同的策略确定了一种产生最佳总体组装的有效方法。我们首先使用单一的k-mer长度组装了大脑和细胞系转录组。接下来,我们测试了组装过程中k-mer长度和覆盖范围的一系列值。最后,我们将来自一系列k值的组装重叠群组合在一起,以生成最终组装。通过比较这些组装结果,我们发现只使用一个k-mer值进行组装并不足以产生良好的组装结果,但将不同k-mer值的重叠群组合起来可以产生更长的重叠群,从而大大提高组装的整体效果。
De novo transcriptome assembly is an important approach in RNA-Seq data analysis and it can help us to reconstruct the transcriptome and investigate gene expression profiles without reference genome sequences. We carried out transcriptome assemblies with two RNA-Seq datasets generated from human brain and cell line, respectively. We then determined an efficient way to yield an optimal overall assembly using three different strategies. We first assembled brain and cell line transcriptome using a single k-mer length. Next we tested a range of values of k-mer length and coverage cutoff in assembling. Lastly, we combined the assembled contigs from a range of k values to generate a final assembly. By comparing these assembly results, we found that using only one k-mer value for assembly is not enough to generate good assembly results, but combining the contigs from different k-mer values could yield longer contigs and greatly improve the overall assembly.