Pan-genome sequence analysis using Panseq: an online tool for the rapid analysis of core and accessory genomic regions.

Pan-genome sequence analysis using Panseq: an online tool for the rapid analysis of core and accessory genomic regions.
复制标题

DOI:
10.1186/1471-2105-11-461
复制
发表时间:
2010-09-15
期刊:
影响因子:
3
通讯作者:
Gannon VP
Gannon VP
中科院分区:
生物学4区
文献类型:
--
作者:
Laing C;Buchanan C;Taboada EN;Zhang Y;Kropinski A;Villegas A;Thomas JE;Gannon VP

文献摘要

参考文献

被引文献

相似文献

细菌物种的泛基因组由核心基因库和辅助基因库组成。辅助基因组被认为是细菌种群遗传变异的重要来源,并通过横向基因转移获得,使细菌亚群更好地适应特定的生态位。低成本和高通量测序平台已经创造了基因组序列数据的指数增长,并为研究许多细菌物种的泛基因组提供了机会。在这项研究中,我们描述了一个新的在线泛基因组序列分析程序,Panseq。Panseq用于鉴定大肠杆菌O 157:H7和E. coli K-12基因组岛。在60个E. coliO 157:H7菌株,经Panseq分析鉴定出65个辅助基因组区域。本研究对6个L.用Panseq提取单核细胞增多性菌株,并分层聚类和可视化。核苷酸核心和二进制附件数据也被用来构建最大简约(MP)树,这是比较多位点序列分型(MLST)生成的MP树。附件和核心树的拓扑结构是相同的,但不同的树使用7个MLST基因座。与穷尽搜索所有可能的组合所需的449 s相比,基因座模块在1 s内发现了10个菌株中100个基因座组中4个基因座的最可变和最具区分性的组合;它还发现了96个基因座中最具区分性的20个基因座。coli O 157:H7 SNP数据集。Panseq根据用户定义的参数确定基因组序列集合中的核心和附属区域。它可以轻松提取基因组或基因组特有的区域,识别共享核心基因组区域内的SNP,根据核心区域内附属区域和SNP的存在/不存在构建用于系统发育程序的文件,并生成图形概述输出。Panseq还包括基因座选择器,其计算辅助基因座或核心基因SNP组中最可变和最具歧视性的基因座。Panseq可以在http://76.70.11.198/panseq上免费获得。Panseq是用Perl编写的。
The pan-genome of a bacterial species consists of a core and an accessory gene pool. The accessory genome is thought to be an important source of genetic variability in bacterial populations and is gained through lateral gene transfer, allowing subpopulations of bacteria to better adapt to specific niches. Low-cost and high-throughput sequencing platforms have created an exponential increase in genome sequence data and an opportunity to study the pan-genomes of many bacterial species. In this study, we describe a new online pan-genome sequence analysis program, Panseq. Panseq was used to identify Escherichia coli O157:H7 and E. coli K-12 genomic islands. Within a population of 60 E. coli O157:H7 strains, the existence of 65 accessory genomic regions identified by Panseq analysis was confirmed by PCR. The accessory genome and binary presence/absence data, and core genome and single nucleotide polymorphisms (SNPs) of six L. monocytogenes strains were extracted with Panseq and hierarchically clustered and visualized. The nucleotide core and binary accessory data were also used to construct maximum parsimony (MP) trees, which were compared to the MP tree generated by multi-locus sequence typing (MLST). The topology of the accessory and core trees was identical but differed from the tree produced using seven MLST loci. The Loci Selector module found the most variable and discriminatory combinations of four loci within a 100 loci set among 10 strains in 1 s, compared to the 449 s required to exhaustively search for all possible combinations; it also found the most discriminatory 20 loci from a 96 loci E. coli O157:H7 SNP dataset. Panseq determines the core and accessory regions among a collection of genomic sequences based on user-defined parameters. It readily extracts regions unique to a genome or group of genomes, identifies SNPs within shared core genomic regions, constructs files for use in phylogeny programs based on both the presence/absence of accessory regions and SNPs within core regions and produces a graphical overview of the output. Panseq also includes a loci selector that calculates the most variable and discriminatory loci among sets of accessory loci or core gene SNPs. Panseq is freely available online at http://76.70.11.198/panseq. Panseq is written in Perl.
DOI: 10.1186/1471-2105-11-142
发表时间: 2010-03-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Kryukov K;Saitou N
通讯作者: Saitou N
DOI: 10.1099/00221287-147-10-2643
发表时间: 2001-10-01
期刊: MICROBIOLOGY-SGM
影响因子: 2.8
作者:
Chetouani, F;Glaser, P;Kunst, F
通讯作者: Kunst, F
DOI: 10.1128/jb.187.16.5537-5551.2005
发表时间: 2005-08-01
影响因子: 3.2
作者:
Nightingale, KK;Windham, K;Wiedmann, M
通讯作者: Wiedmann, M
DOI: 10.1099/jmm.0.45971-0
发表时间: 2005-10-01
影响因子: 3
作者:
Best, EL;Fox, AJ;Bolton, FJ
通讯作者: Bolton, FJ
DOI: 10.1101/gr.078212.108
发表时间: 2008-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Heng;Ruan, Jue;Durbin, Richard
通讯作者: Durbin, Richard