Comparison of different sequencing and assembly strategies for a repeat-rich fungal genome, Ophiocordyceps sinensis

Comparison of different sequencing and assembly strategies for a repeat-rich fungal genome, Ophiocordyceps sinensis
复制标题

富含重复的真菌基因组冬虫夏草的不同测序和组装策略的比较

DOI:
10.1016/j.mimet.2016.06.025
复制
发表时间:
2016-09-01
影响因子:
2.2
通讯作者:
Yao, Yi-Jian
Yao, Yi-Jian
中科院分区:
生物学4区
文献类型:
--
作者:
Li, Yi;Hsiang, Tom;Yao, Yi-Jian

文献摘要

被引文献

相似文献

冬虫夏草(Ophiocordyceps sinensis)是世界上最昂贵的药用真菌之一,几个世纪以来一直被用作中药。在最近的一份报告中,在Roche 454 (223 Mb)和Illumina HiSeq (10.6 Gb)测序数据组装后,发现该真菌的基因组被大量重复元件扩增,产生87.7 Mb的基因组,N50支架长度为12 kb,预测基因6972个。为了测试是否可以通过更深入的测序来改善组装,并评估最佳组装所需的数据量,在Illumina HiSeq平台上对单个子囊孢子分离物(菌株1229)的基因组DNA提取进行了多次基因组测序(总数据为25 Gb)。程序集是使用不同的数据类型(原始的还是修剪过的)和数据量生成的,并使用三个免费的汇编程序(ABySS、SOAP和Velvet)。在几乎所有情况下,与未修剪的数据相比,为低质量的基调用修剪数据并不会为程序集提供更高的N50值,并且增加输入数据量(即序列读取)并不总是会导致更高的N50值。根据汇编程序和数据类型的不同,最大N50在总读取数据的50%到90%之间达到,相当于100倍到200倍的覆盖率。与先前发表的版本相比,该基因组组装草图得到了改进,得到了114 Mb的组装,脚手架N50为70 kb,预测基因为9610。在本研究预测的基因中,通过RNA-Seq分析验证了9213个基因,其中8896个为单子基因。来自基因组和转录组分析的证据表明,物种组装可以通过确定的输入材料(例如单倍体单子囊孢子分离物)来改善,而不需要多种测序技术,多种文库大小或低质量碱基调用的数据修剪,基因组覆盖率在100到200倍之间。(C) 2016 Elsevier B.V.版权所有
Ophiocordyceps sinensis is one of the most expensive medicinal fungi world-wide, and has been used as a traditional Chinese medicine for centuries. In a recent report, the genome of this fungus was found to be expanded by extensive repetitive elements after assembly of Roche 454 (223 Mb) and Illumina HiSeq (10.6 Gb) sequencing data, producing a genome of 87.7 Mb with an N50 scaffold length of 12 kb and 6972 predicted genes. To test whether the assembly could be improved by deeper sequencing and to assess the amount of data needed for optimal assembly, genomic sequencing was run several times on genomic DNA extractions of a single ascospore isolate (strain 1229) on an Illumina HiSeq platform (25 Gb total data). Assemblies were produced using different data types (raw vs. trimmed) and data amounts, and using three freely available assembly programs (ABySS, SOAP and Velvet). In nearly all cases, trimming the data for low quality base calls did not provide assemblies with higher N50 values compared to the non-trimmed data, and increasing the amount of input data (i.e. sequence reads) did not always lead to higher N50 values. Depending on the assembly program and data type, the maximal N50 was reached with between 50% to 90% of the total read data, equivalent to 100 x to 200 x coverage. The draft genome assembly was improved over the previously published version resulting in a 114 Mb assembly, scaffold N50 of 70 kb and 9610 predicted genes. Among the predicted genes, 9213 were validated by RNA-Seq analysis in this study, of which 8896 were found to be singletons. Evidence from genome and transcriptome analyses indicated that species assemblies could be improved with defined input material (e.g. haploid mono-ascospore isolate) without the requirement of multiple sequencing technologies, multiple library sizes or data trimming for low quality base calls, and with genome coverages between 100x and 200x. (C) 2016 Elsevier B.V. All rights reserved.