Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications.

Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications.
复制标题

开放式细菌种群基因组学:BIGSDB软件,pubmlst.org网站及其应用。

DOI:
10.12688/wellcomeopenres.14826.1
复制
发表时间:
2018
影响因子:
--
通讯作者:
Maiden MCJ
Maiden MCJ
中科院分区:
其他
文献类型:
--
作者:
Jolley KA;Bray JE;Maiden MCJ

文献摘要

被引文献

相似文献

PubMLST.org网站拥有一系列开放访问的经过管理的数据库,这些数据库将种群序列数据与100多个不同微生物物种和属的种源和表型信息整合在一起。尽管PubMLST网站是1998年第一个多位点序列分型(MLST)方案开发的一部分,但它使用的软件-细菌分离基因组序列数据库(BIGSdb)-使PubMLST能够包括所有级别的序列数据,从单基因序列到包括完整的、完成的基因组。在这里,我们描述了从出版到2018年6月BIGSdb软件的发展,并展示了该平台如何在广泛的应用中实现微生物种群基因组学。该系统基于对微生物基因组的逐个基因分析,对每个存放的序列进行注释和整理,以识别存在的基因并系统地编目它们的变异。最初的目的是作为一种用分型方案表征分离物的手段,合成具有来源和表型数据的遗传变异的序列和记录允许高度可扩展的(数万个分离物的全基因组序列数据)手段来解决广泛的功能性问题,包括:预测抗菌素耐药性;可能与疫苗抗原的交叉反应;以及导致关键表型的不同变体的功能活动。*可以包括的序列、遗传位点、等位基因变体或方案(位点组合)的数量没有限制,使每个数据库能够代表相关人群的遗传变异的不断扩大的目录。除了提供网络可访问的分析和指向第三方分析和可视化工具的链接外,BIGSdb软件还包括REST风格的应用程序编程接口(API),使您能够访问第三方应用程序和数据分析管道的所有基础数据。
The PubMLST.org website hosts a collection of open-access, curated databases that integrate population sequence data with provenance and phenotype information for over 100 different microbial species and genera.  Although the PubMLST website was conceived as part of the development of the first multi-locus sequence typing (MLST) scheme in 1998 the software it uses, the Bacterial Isolate Genome Sequence database (BIGSdb, published in 2010), enables PubMLST to include all levels of sequence data, from single gene sequences up to and including complete, finished genomes.  Here we describe developments in the BIGSdb software made from publication to June 2018 and show how the platform realises microbial population genomics for a wide range of applications.  The system is based on the gene-by-gene analysis of microbial genomes, with each deposited sequence annotated and curated to identify the genes present and systematically catalogue their variation.  Originally intended as a means of characterising isolates with typing schemes, the synthesis of sequences and records of genetic variation with provenance and phenotype data permits highly scalable (whole genome sequence data for tens of thousands of isolates) means of addressing a wide range of functional questions, including: the prediction of antimicrobial resistance; likely cross-reactivity with vaccine antigens; and the functional activities of different variants that lead to key phenotypes.  There are no limitations to the number of sequences, genetic loci, allelic variants or schemes (combinations of loci) that can be included, enabling each database to represent an expanding catalogue of the genetic variation of the population in question.  In addition to providing web-accessible analyses and links to third-party analysis and visualisation tools, the BIGSdb software includes a RESTful application programming interface (API) that enables access to all the underlying data for third-party applications and data analysis pipelines.