De novo assembly of the Pseudomonas syringae pv. syringae B728a genome using Illumina/Solexa short sequence reads

De novo assembly of the Pseudomonas syringae pv. syringae B728a genome using Illumina/Solexa short sequence reads
复制标题

DOI:
10.1111/j.1574-6968.2008.01441.x
复制
发表时间:
2009-02-01
影响因子:
2.1
通讯作者:
Studholme, David J.
Studholme, David J.
中科院分区:
生物学4区
文献类型:
--
作者:
Farrer, Rhys A.;Kemen, Eric;Studholme, David J.

文献摘要

被引文献

相似文献

Illumina 的基因组分析仪可生成超短序列读数,长度通常为 36 个核苷酸,主要用于重新测序。我们测试了该技术在丁香假单胞菌 pv 的 6 Mbp 基因组上从头序列组装的潜力。 syringae B728a 具有多个免费可用的组装软件包。使用未配对的数据集,velvet 将 > 96% 的基因组组装成重叠群,N50 长度为 8289 个核苷酸,错误率为 0.33%。 edena 生成更小的重叠群(N50 为 4192 个核苷酸)和相当的错误率。 ssake 和 vcake 产生的重叠群较短,错误率非常高。携带 400 bp 插入片段的配对末端序列数据的组装产生了更长的重叠群(N50 高达 15 628 个核苷酸),但错误率增加(0.5%)。重叠群长度和错误率对参数值的选择非常敏感。非编码 RNA 基因在从头组装中解析度较差,而超过 90% 的蛋白质编码基因在其全长上的组装精度为 100%。这项研究表明,在实践中,36 核苷酸读数的从头组装可以从约 40 倍深度的序列数据集中生成相当准确的组装。这些草图组件可用于以非常经济的低成本探索生物体的蛋白质组潜力。
Illumina's Genome Analyzer generates ultra-short sequence reads, typically 36 nucleotides in length, and is primarily intended for resequencing. We tested the potential of this technology for de novo sequence assembly on the 6 Mbp genome of Pseudomonas syringae pv. syringae B728a with several freely available assembly software packages. Using an unpaired data set, velvet assembled > 96% of the genome into contigs with an N50 length of 8289 nucleotides and an error rate of 0.33%. edena generated smaller contigs (N50 was 4192 nucleotides) and comparable error rates. ssake and vcakeyielded shorter contigs with very high error rates. Assembly of paired-end sequence data carrying 400 bp inserts produced longer contigs (N50 up to 15 628 nucleotides), but with increased error rates (0.5%). Contig length and error rate were very sensitive to the choice of parameter values. Noncoding RNA genes were poorly resolved in de novo assemblies, while > 90% of the protein-coding genes were assembled with 100% accuracy over their full length. This study demonstrates that, in practice, de novo assembly of 36-nucleotide reads can generate reasonably accurate assemblies from about 40 x deep sequence data sets. These draft assemblies are useful for exploring an organism's proteomic potential, at a very economic low cost.