Comparison of assembly algorithms for improving rate of metatranscriptomic functional annotation.

Comparison of assembly algorithms for improving rate of metatranscriptomic functional annotation.
复制标题

DOI:
10.1186/2049-2618-2-39
复制
发表时间:
2014
期刊:
影响因子:
15.5
通讯作者:
Parkinson J
Parkinson J
中科院分区:
生物学1区
文献类型:
--
作者:
Celaj A;Markle J;Danska J;Parkinson J

文献摘要

参考文献

被引文献

相似文献

通过高通量RNA测序(“元转录组学”)进行的微生物组范围的基因表达谱分析提供了功能性地询问复杂微生物群落的有力手段。成功利用这些数据集的关键是能够自信地将相对较短的序列读数与已知的细菌转录物相匹配。在不存在参考基因组的情况下,在数据库搜索策略之前,可以通过将读段组装成更长的连续序列(“重叠群”)来增强这种注释工作。由于来自同源转录物的读段可能来源于以不同丰度水平表示的几个物种,因此目前尚不清楚当前组装管道对于元转录组数据集的表现如何。在这里,我们评估了四个目前采用的汇编程序,包括从头转录组汇编程序- Trinity和Oases;宏基因组汇编程序-Metavelet;和最近开发的元转录组汇编程序IDBA-MT的性能。我们评估了组装器在先前发表的来自1型糖尿病近交系非肥胖糖尿病小鼠模型大肠的单端RNA序列读数数据集上的性能。我们发现Trinity表现最好,这是通过组装的重叠群、分配给重叠群的读段和可以注释到已知细菌转录物的读段的数量来判断的。只有15.5%的RNA序列读段可以注释到已知的转录本,而Trinity组装的结果为50.3%。从相同的小鼠样品产生的双端读段导致适度的性能增益。数据库搜索估计,组装不太可能错误地合并共享相似区域的多个不相关基因(<2%的重叠群)。基于10个物种的模拟数据集证实了这些发现。基于72个物种的更复杂的模拟数据集发现,引入了比测序质量预期更大的组装错误。通过对组装性能的详细评估,本研究提供的见解将有助于推动未来元转录组分析的设计。元转录组数据集的组装极大地改进了读段注释。在评估的四个汇编器中,Trinity提供了最好的性能。对于更复杂的数据集,从共享相当大的序列相似性的转录物生成的读段可能是显著组装错误的来源,这表明需要在组装之前基于共同的分类学来源整理读段。
Microbiome-wide gene expression profiling through high-throughput RNA sequencing (‘metatranscriptomics’) offers a powerful means to functionally interrogate complex microbial communities. Key to successful exploitation of these datasets is the ability to confidently match relatively short sequence reads to known bacterial transcripts. In the absence of reference genomes, such annotation efforts may be enhanced by assembling reads into longer contiguous sequences (‘contigs’), prior to database search strategies. Since reads from homologous transcripts may derive from several species, represented at different abundance levels, it is not clear how well current assembly pipelines perform for metatranscriptomic datasets. Here we evaluate the performance of four currently employed assemblers including de novo transcriptome assemblers - Trinity and Oases; the metagenomic assembler - Metavelvet; and the recently developed metatranscriptomic assembler IDBA-MT. We evaluated the performance of the assemblers on a previously published dataset of single-end RNA sequence reads derived from the large intestine of an inbred non-obese diabetic mouse model of type 1 diabetes. We found that Trinity performed best as judged by contigs assembled, reads assigned to contigs, and number of reads that could be annotated to a known bacterial transcript. Only 15.5% of RNA sequence reads could be annotated to a known transcript in contrast to 50.3% with Trinity assembly. Paired-end reads generated from the same mouse samples resulted in modest performance gains. A database search estimated that the assemblies are unlikely to erroneously merge multiple unrelated genes sharing a region of similarity (<2% of contigs). A simulated dataset based on ten species confirmed these findings. A more complex simulated dataset based on 72 species found that greater assembly errors were introduced than is expected by sequencing quality. Through the detailed evaluation of assembly performance, the insights provided by this study will help drive the design of future metatranscriptomic analyses. Assembly of metatranscriptome datasets greatly improved read annotation. Of the four assemblers evaluated, Trinity provided the best performance. For more complex datasets, reads generated from transcripts sharing considerable sequence similarity can be a source of significant assembly error, suggesting a need to collate reads on the basis of common taxonomic origin prior to assembly.
DOI: 10.1093/nar/gkt1196
发表时间: 2014-01
影响因子: 14.9
作者:
Flicek P;Amode MR;Barrell D;Beal K;Billis K;Brent S;Carvalho-Silva D;Clapham P;Coates G;Fitzgerald S;Gil L;Girón CG;Gordon L;Hourlier T;Hunt S;Johnson N;Juettemann T;Kähäri AK;Keenan S;Kulesha E;Martin FJ;Maurel T;McLaren WM;Murphy DN;Nag R;Overduin B;Pignatelli M;Pritchard B;Pritchard E;Riat HS;Ruffier M;Sheppard D;Taylor K;Thormann A;Trevanion SJ;Vullo A;Wilder SP;Wilson M;Zadissa A;Aken BL;Birney E;Cunningham F;Harrow J;Herrero J;Hubbard TJ;Kinsella R;Muffato M;Parker A;Spudich G;Yates A;Zerbino DR;Searle SM
通讯作者: Searle SM
DOI: 10.1126/science.1241214
发表时间: 2013-09-06
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Ridaura VK;Faith JJ;Rey FE;Cheng J;Duncan AE;Kau AL;Griffin NW;Lombard V;Henrissat B;Bain JR;Muehlbauer MJ;Ilkayeva O;Semenkovich CF;Funai K;Hayashi DK;Lyle BJ;Martini MC;Ursell LK;Clemente JC;Van Treuren W;Walters WA;Knight R;Newgard CB;Heath AC;Gordon JI
通讯作者: Gordon JI
DOI: 10.1093/bioinformatics/bts094
发表时间: 2012-04-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Schulz MH;Zerbino DR;Vingron M;Birney E
通讯作者: Birney E
DOI: 10.1126/science.1233521
发表时间: 2013-03-01
期刊: SCIENCE
影响因子: 56.9
作者:
Markle, Janet G. M.;Frank, Daniel N.;Danska, Jayne S.
通讯作者: Danska, Jayne S.
元基因组,荟萃分析组和单细胞测序揭示了对深水地平线溢油的微生物反应。
DOI: 10.1038/ismej.2012.59
发表时间: 2012-09
期刊: The ISME journal
影响因子: --
作者:
通讯作者: --