Next generation sequencing data of a defined microbial mock community.

Next generation sequencing data of a defined microbial mock community.
复制标题

DOI:
10.1038/sdata.2016.81
复制
发表时间:
2016-09-27
期刊:
影响因子:
9.8
通讯作者:
Woyke T
Woyke T
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Singer E;Andreopoulos B;Bowers RM;Lee J;Deshpande S;Chiniquy J;Ciobanu D;Klenk HP;Zane M;Daum C;Clum A;Cheng JF;Copeland A;Woyke T

文献摘要

被引文献

相似文献

生成由具有完整参考基因组的生物体组成的确定群落的序列数据对于新的基因组序列分析方法(包括组装和分箱工具)的基准测试是必不可少的。此外,新测序文库方案和平台的验证以评估关键组分,如测序错误和偏差,依赖于此类数据集。我们在这里报告了下一代宏基因组序列数据的一个定义的模拟社区(模拟细菌ARchaea社区; MBARC-26),由23个细菌和3个古细菌菌株完成的基因组。这些菌株跨越10门和14类,GC含量,基因组大小,重复内容的范围,并涵盖了不同的丰度谱。描述了该模拟群落的短读段Illumina和长读段PacBio SMRT序列。这些数据为科学界提供了宝贵的资源,使生物信息学工具能够进行广泛的基准测试和比较评估,而无需模拟数据。因此,这些数据可以帮助改进我们目前的序列数据分析工具包,并激发开发新工具的兴趣。
Generating sequence data of a defined community composed of organisms with complete reference genomes is indispensable for the benchmarking of new genome sequence analysis methods, including assembly and binning tools. Moreover the validation of new sequencing library protocols and platforms to assess critical components such as sequencing errors and biases relies on such datasets. We here report the next generation metagenomic sequence data of a defined mock community (Mock Bacteria ARchaea Community; MBARC-26), composed of 23 bacterial and 3 archaeal strains with finished genomes. These strains span 10 phyla and 14 classes, a range of GC contents, genome sizes, repeat content and encompass a diverse abundance profile. Short read Illumina and long-read PacBio SMRT sequences of this mock community are described. These data represent a valuable resource for the scientific community, enabling extensive benchmarking and comparative evaluation of bioinformatics tools without the need to simulate data. As such, these data can aid in improving our current sequence data analysis toolkit and spur interest in the development of new tools.