RNA sequencing read depth requirement for optimal transcriptome coverage in Hevea brasiliensis.

RNA sequencing read depth requirement for optimal transcriptome coverage in Hevea brasiliensis.
复制标题

DOI:
10.1186/1756-0500-7-69
复制
发表时间:
2014-02-01
期刊:
影响因子:
1.8
通讯作者:
Mohd-Zainuddin Z
Mohd-Zainuddin Z
中科院分区:
其他
文献类型:
--
作者:
Chow KS;Ghazali AK;Hoh CC;Mohd-Zainuddin Z

文献摘要

被引文献

相似文献

组装从头转录组的关注点之一是确定确保全面覆盖特定样品中表达的基因所需的读取序列的量。在这份报告中,我们描述了使用Illumina配对末端RNA-Seq(PE RNA-Seq)读取橡胶树(橡胶树)树皮设计一个转录本映射方法,用于估计深度转录组覆盖所需的读取量。我们使用Oases组装器在一系列k-mer大小上基于16 Gb Illumina PE RNA-Seq读段优化了橡胶树树皮转录组的组装。然后,我们基于转录本N50长度和转录本作图统计评估组装质量,所述转录本作图统计与(a)具有完整开放阅读框的已知橡胶树cDNA,(B)一组核心真核基因和(c)橡胶树基因组支架有关。随后是系统的转录物作图过程,其中将来自一系列增量的树皮转录物的子组装体与来自整个树皮转录组组装体的转录物进行比对。该练习用于将读取量与转录本映射水平的程度联系起来,后者是样本中表达的基因转录本覆盖率的指标。随着读取量或数据大小增加到16 Gb,映射到整个树皮组件的转录本的数量接近饱和。随后生成颜色矩阵以说明与总样品转录物的覆盖程度相关的测序深度要求。我们设计了一个程序,“转录映射饱和测试”,以估计深度覆盖转录组所需的RNA-Seq读数的量。对于橡胶树从头组装,我们建议产生5-8 Gb之间的读取,其中约90%的转录本覆盖率可以用优化的k-mer和转录本N50长度来实现。该方法背后的原理也可以应用于其他非模型植物,或来自其他第二代测序平台的读数。
One of the concerns of assembling de novo transcriptomes is determining the amount of read sequences required to ensure a comprehensive coverage of genes expressed in a particular sample. In this report, we describe the use of Illumina paired-end RNA-Seq (PE RNA-Seq) reads from Hevea brasiliensis (rubber tree) bark to devise a transcript mapping approach for the estimation of the read amount needed for deep transcriptome coverage. We optimized the assembly of a Hevea bark transcriptome based on 16 Gb Illumina PE RNA-Seq reads using the Oases assembler across a range of k-mer sizes. We then assessed assembly quality based on transcript N50 length and transcript mapping statistics in relation to (a) known Hevea cDNAs with complete open reading frames, (b) a set of core eukaryotic genes and (c) Hevea genome scaffolds. This was followed by a systematic transcript mapping process where sub-assemblies from a series of incremental amounts of bark transcripts were aligned to transcripts from the entire bark transcriptome assembly. The exercise served to relate read amounts to the degree of transcript mapping level, the latter being an indicator of the coverage of gene transcripts expressed in the sample. As read amounts or datasize increased toward 16 Gb, the number of transcripts mapped to the entire bark assembly approached saturation. A colour matrix was subsequently generated to illustrate sequencing depth requirement in relation to the degree of coverage of total sample transcripts. We devised a procedure, the “transcript mapping saturation test”, to estimate the amount of RNA-Seq reads needed for deep coverage of transcriptomes. For Hevea de novo assembly, we propose generating between 5–8 Gb reads, whereby around 90% transcript coverage could be achieved with optimized k-mers and transcript N50 length. The principle behind this methodology may also be applied to other non-model plants, or with reads from other second generation sequencing platforms.