PANINI: Pangenome Neighbour Identification for Bacterial Populations

PANINI: Pangenome Neighbour Identification for Bacterial Populations
复制标题

DOI:
10.1099/mgen.0.000220
复制
发表时间:
2019-04-01
期刊:
影响因子:
3.9
通讯作者:
Aanensen, David M.
Aanensen, David M.
中科院分区:
生物学2区
文献类型:
--
作者:
Abudahab, Khalil;Prada, Joaquin M.;Aanensen, David M.

文献摘要

被引文献

相似文献

细菌种群进化的基因组分析的标准主力是核心基因组突变的系统发育建模。然而,如果不考虑辅助基因组,则可能丢失大量关于不同种群中进化和传播过程的信息。在这里,我们介绍PANINI(细菌种群的泛基因组邻居识别),这是一种计算可扩展的方法,用于使用基于t-SNE(t分布随机邻居嵌入)算法的随机邻居嵌入的无监督机器学习来识别数据集中每个分离株的邻居。PANINI基于浏览器,并与Microreact平台集成,用于快速在线可视化和探索核心和辅助基因组进化信号,以及相关的流行病学、地理、时间和其他元数据。与单克隆和多克隆肺炎球菌种群的几个案例研究,以证明从基因内容数据识别生物学上重要的信号的能力。PANINI可在http://panini.pathogen.watch上获得,代码在http://gitlab.com/cgps/panini上。
The standard workhorse for genomic analysis of the evolution of bacterial populations is phylogenetic modelling of mutations in the core genome. However, a notable amount of information about evolutionary and transmission processes in diverse populations can be lost unless the accessory genome is also taken into consideration. Here, we introduce PANINI (Pangenome Neighbour Identification for Bacterial Populations), a computationally scalable method for identifying the neighbours for each isolate in a data set using unsupervised machine learning with stochastic neighbour embedding based on the t-SNE (t-distributed stochastic neighbour embedding) algorithm. PANINI is browser-based and integrates with the Microreact platform for rapid online visualization and exploration of both core and accessory genome evolutionary signals, together with relevant epidemiological, geographical, temporal and other metadata. Several case studies with single- and multi-clone pneumococcal populations are presented to demonstrate the ability to identify biologically important signals from gene content data. PANINI is available at http://panini.pathogen.watch and code at http://gitlab.com/cgps/panini.