Optimization of de novo short read assembly of seabuckthorn (Hippophae rhamnoides L.) transcriptome.

Optimization of de novo short read assembly of seabuckthorn (Hippophae rhamnoides L.) transcriptome.
复制标题

DOI:
10.1371/journal.pone.0072516
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Chand Sharma P
Chand Sharma P
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ghangal R;Chaudhary S;Jain M;Purty RS;Chand Sharma P

文献摘要

参考文献

被引文献

相似文献

沙棘(Hippophae rhamnoides L.)自古以来就以其药用、营养和环境重要性而闻名。然而,在描述这种神奇植物的基因组和转录组方面所做的努力非常有限。在这里,我们报告使用下一代大规模并行测序技术(Illumina平台)和从头组装,以获得沙棘转录组的全面视图。我们使用六种组装工具组装了86,253,874个高质量短读段。在我们手中,考虑到各种组装质量参数,发现遵循两步程序的非冗余短读段的组装是最好的。最初,ABySS工具是按照加性k-mer方法使用的。随后将组装的转录物置于TGICL套件中。最后,从头短读组装产生88,297个转录物(> 100 bp),代表约53 Mb的沙棘转录组。平均长度为610 bp,N50长度为1198 BP,91%的短读段唯一地映射回沙棘转录组。共有41,340个(46.8%)转录本与NCBI的nr蛋白质数据库中存在的序列显示出显著的相似性(E值<1 E-06)。我们还筛选了组装的转录本中是否存在转录因子和简单重复序列。我们的策略涉及使用短读汇编程序(ABySS),然后TGICL将是有用的研究人员与非模式生物的转录组在节省时间和降低数据管理的复杂性。本研究所获得的沙棘转录组数据为基因发现和功能分子标记的开发提供了宝贵的资源。
Seabuckthorn ( Hippophae rhamnoides L.) is known for its medicinal, nutritional and environmental importance since ancient times. However, very limited efforts have been made to characterize the genome and transcriptome of this wonder plant. Here, we report the use of next generation massive parallel sequencing technology (Illumina platform) and de novo assembly to gain a comprehensive view of the seabuckthorn transcriptome. We assembled 86,253,874 high quality short reads using six assembly tools. At our hand, assembly of non-redundant short reads following a two-step procedure was found to be the best considering various assembly quality parameters. Initially, ABySS tool was used following an additive k-mer approach. The assembled transcripts were subsequently subjected to TGICL suite. Finally, de novo short read assembly yielded 88,297 transcripts (> 100 bp), representing about 53 Mb of seabuckthorn transcriptome. The average length of transcripts was 610 bp, N50 length 1198 BP and 91% of the short reads uniquely mapped back to seabuckthorn transcriptome. A total of 41,340 (46.8%) transcripts showed significant similarity with sequences present in nr protein databases of NCBI (E-value < 1E-06). We also screened the assembled transcripts for the presence of transcription factors and simple sequence repeats. Our strategy involving the use of short read assembler (ABySS) followed by TGICL will be useful for the researchers working with a non-model organism’s transcriptome in terms of saving time and reducing complexity in data management. The seabuckthorn transcriptome data generated here provide a valuable resource for gene discovery and development of functional molecular markers.
DOI: 10.1093/bioinformatics/bts094
发表时间: 2012-04-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Schulz MH;Zerbino DR;Vingron M;Birney E
通讯作者: Birney E
DOI: 10.1101/gr.097261.109
发表时间: 2010-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Ruiqiang;Zhu, Hongmei;Wang, Jun
通讯作者: Wang, Jun
DOI: 10.1016/j.jep.2004.02.016
发表时间: 2004-06-01
影响因子: 5.4
作者:
Sezik, E;Yesilada, E;Honda, G
通讯作者: Honda, G
DOI: 10.1186/1471-2105-8-42
发表时间: 2007-02-07
期刊: BMC bioinformatics
影响因子: 3
作者:
Riaño-Pachón DM;Ruzicic S;Dreyer I;Mueller-Roeber B
通讯作者: Mueller-Roeber B
DOI: 10.1016/s0378-8741(02)00333-1
发表时间: 2003-02-01
影响因子: 5.4
作者:
Shinwari, ZK;Gilani, SS
通讯作者: Gilani, SS