Short read Illumina data for the de novo assembly of a non-model snail species transcriptome (Radix balthica, Basommatophora, Pulmonata), and a comparison of assembler performance.

Short read Illumina data for the de novo assembly of a non-model snail species transcriptome (Radix balthica, Basommatophora, Pulmonata), and a comparison of assembler performance.
复制标题

DOI:
10.1186/1471-2164-12-317
复制
发表时间:
2011-06-16
期刊:
影响因子:
4.4
通讯作者:
Pfenninger M
Pfenninger M
中科院分区:
生物学2区
文献类型:
--
作者:
Feldmeyer B;Wheat CW;Krezdorn N;Rotter B;Pfenninger M

文献摘要

参考文献

被引文献

相似文献

直到最近,Solexa/Illumina系统上的阅读长度太短,无法在没有参考序列的情况下可靠地组装转录本,特别是对于非模式生物。然而,随着当前版本中可获得高达100个核苷酸的读取长度,没有参考基因组的组装应该是可能的。在这项研究中,我们通过Illumina对标准化转录组的测序,创建了常见池塘螺根的EST数据集。从重叠群的数量、长度、覆盖深度、在各种BLAST搜索中的质量以及与线粒体基因的比对等方面比较了三种不同的短读汇编器的性能。对归一化RNA池的单次测序运行产生了16,923,850对末端读取,读取长度的中位数为61个碱基。天鹅绒、绿洲和SeqMan ngen产生的组件在重叠群总数、重叠群长度、通过对各种数据库进行BLAST搜索获得的基因命中的数量和质量以及在mt基因组比较中的重叠群性能方面存在差异。虽然天鹅绒产生的重叠群总数最高,但其中很大一部分重叠群大小(200个基点),并在BLAST搜索和mt基因组比对中提供多余的匹配。最好的整体重叠群性能来自于ngen组装。它产生了第二大数量的重叠群,平均而言与绿洲重叠群相当,但在对不同参考数据库的四次BLAST搜索中,有两次给出了最高的基因命中率。随后的四个重叠群集合的元组装导致更大的重叠群、更少的冗余和更高的BLAST命中数量。我们的结果使用Illumina测序数据证明了第一个非模式物种的从头转录组组装。我们表明,使用这种方法的从头转录组组装产生了对下游应用有用的结果,特别是如果重叠群集的元组装被用于提高重叠群质量。这些结果突显了组装方法持续改进的必要性。
Until recently, read lengths on the Solexa/Illumina system were too short to reliably assemble transcriptomes without a reference sequence, especially for non-model organisms. However, with read lengths up to 100 nucleotides available in the current version, an assembly without reference genome should be possible. For this study we created an EST data set for the common pond snail Radix balthica by Illumina sequencing of a normalized transcriptome. Performance of three different short read assemblers was compared with respect to: the number of contigs, their length, depth of coverage, their quality in various BLAST searches and the alignment to mitochondrial genes. A single sequencing run of a normalized RNA pool resulted in 16,923,850 paired end reads with median read length of 61 bases. The assemblies generated by VELVET, OASES, and SeqMan NGEN differed in the total number of contigs, contig length, the number and quality of gene hits obtained by BLAST searches against various databases, and contig performance in the mt genome comparison. While VELVET produced the highest overall number of contigs, a large fraction of these were of small size (< 200bp), and gave redundant hits in BLAST searches and the mt genome alignment. The best overall contig performance resulted from the NGEN assembly. It produced the second largest number of contigs, which on average were comparable to the OASES contigs but gave the highest number of gene hits in two out of four BLAST searches against different reference databases. A subsequent meta-assembly of the four contig sets resulted in larger contigs, less redundancy and a higher number of BLAST hits. Our results document the first de novo transcriptome assembly of a non-model species using Illumina sequencing data. We show that de novo transcriptome assembly using this approach yields results useful for downstream applications, in particular if a meta-assembly of contig sets is used to increase contig quality. These results highlight the ongoing need for improvements in assembly methodology.
DOI: 10.1186/1471-2164-10-345
发表时间: 2009-07-31
期刊: BMC genomics
影响因子: 4.4
作者:
Kristiansson E;Asker N;Förlin L;Larsson DG
通讯作者: Larsson DG
DOI: 10.1186/1471-2164-11-180
发表时间: 2010-03-16
期刊: BMC genomics
影响因子: 4.4
作者:
Parchman TL;Geist KS;Grahnen JA;Benkman CW;Buerkle CA
通讯作者: Buerkle CA
DOI: 10.1186/1471-2164-10-234
发表时间: 2009-05-19
期刊: BMC genomics
影响因子: 4.4
作者:
Hahn DA;Ragland GJ;Shoemaker DD;Denlinger DL
通讯作者: Denlinger DL
DOI: 10.1093/molbev/msp188
发表时间: 2009-12-01
影响因子: 10.7
作者:
Gibbons, John G.;Janson, Eric M.;Rokas, Antonis
通讯作者: Rokas, Antonis
DOI: 10.1101/gr.5145806
发表时间: 2007-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Emrich, Scott J.;Barbazuk, W. Brad;Schnable, Patrick S.
通讯作者: Schnable, Patrick S.