A curated dataset of complete Enterobacteriaceae plasmids compiled from the NCBI nucleotide database.

A curated dataset of complete Enterobacteriaceae plasmids compiled from the NCBI nucleotide database.
复制标题

DOI:
10.1016/j.dib.2017.04.024
复制
发表时间:
2017-06
期刊:
影响因子:
1.2
通讯作者:
Stoesser N
Stoesser N
中科院分区:
其他
文献类型:
--
作者:
Orlek A;Phan H;Sheppard AE;Doumith M;Ellington M;Peto T;Crook D;Walker AS;Woodford N;Anjum MF;Stoesser N

文献摘要

被引文献

相似文献

目前在NCBI核苷酸数据库中公开提供了数千个质粒序列,但它们没有可靠的注释来区分完整的质粒和质粒片段,如基因或结构序列;因此,获取完整的质粒用于下游分析是具有挑战性的。在这里,我们提出了一个来自临床相关肠杆菌科的完整细菌质粒的策划数据集。该数据集是从NCBI核苷酸数据库中编译而来,使用的管理步骤旨在排除不完整的质粒序列和被错误注释为质粒的染色体序列。超过2000个完整的质粒序列包含在精心策划的质粒数据集中。还提供了翻译所有6帧中的每个完整质粒核苷酸序列产生的蛋白质序列。随附的研究文章“订购暴徒:对复制子和暴徒类型的见解……”(Orlek et al., 2017)[1]中提出了对数据集的进一步分析和讨论。整理的质粒序列在Figshare存储库中公开可用。
Thousands of plasmid sequences are now publicly available in the NCBI nucleotide database, but they are not reliably annotated to distinguish complete plasmids from plasmid fragments, such as gene or contig sequences; therefore, retrieving complete plasmids for downstream analyses is challenging. Here we present a curated dataset of complete bacterial plasmids from the clinically relevant Enterobacteriaceae family. The dataset was compiled from the NCBI nucleotide database using curation steps designed to exclude incomplete plasmid sequences, and chromosomal sequences misannotated as plasmids. Over 2000 complete plasmid sequences are included in the curated plasmid dataset. Protein sequences produced from translating each complete plasmid nucleotide sequence in all 6 frames are also provided. Further analysis and discussion of the dataset is presented in an accompanying research article: “Ordering the mob: insights into replicon and MOB typing…” (Orlek et al., 2017) [1]. The curated plasmid sequences are publicly available in the Figshare repository.