snpTree--a web-server to identify and construct SNP trees from whole genome sequence data.

snpTree--a web-server to identify and construct SNP trees from whole genome sequence data.
复制标题

DOI:
10.1186/1471-2164-13-s7-s6
复制
发表时间:
2012
期刊:
影响因子:
4.4
通讯作者:
Aarestrup FM
Aarestrup FM
中科院分区:
生物学2区
文献类型:
--
作者:
Leekitcharoenphon P;Kaas RS;Thomsen MC;Friis C;Rasmussen S;Aarestrup FM

文献摘要

被引文献

相似文献

全基因组测序(WGS)的进步和不断降低的经济成本,将很快使这项技术可用于常规的传染病流行病学。在流行病学研究中,爆发分离株的多样性非常小,需要进行广泛的基因组分析来区分和分类分离株。单核苷酸多态性(SNPs)分析是目前应用最广泛的方法之一。目前,有不同的工具和方法来识别SNP,包括各种选项和截止值。此外,目前所有的方法都需要生物信息学技能。因此,我们缺乏一个标准和简单的自动工具来确定SNPs和构建系统发育树从WGS数据。在这里,我们介绍snpTree,一个在线自动SNP分析服务器。该工具由不同的SNP分析套件、Perl和Python脚本组成。snpTree可以从WGS以及从组装的基因组或重叠群鉴定SNP并构建系统发育树。fastq格式的WGS数据通过BWA与参考基因组比对,而fasta格式的重叠群通过Nucmer处理。基于参考基因组上的位置连接SNP,并使用FastTree和perl脚本从连接的SNP构建树。在线服务器采用HTML、Java和python脚本实现。使用四个公开的细菌WGS数据集(霍乱弧菌、沙门氏菌、沙门氏菌)评估服务器。aureus CC398、S.鼠伤寒和M.肺结核)。前三种情况的评估结果对于原始读数和组装的基因组是一致的。在后一种情况下,原始出版物涉及对SNP的广泛过滤,这不能使用snpTree重复。snpTree服务器是流行病学研究中快速标准化和自动SNP分析的一个易于使用的选项,也适用于生物信息学经验有限的用户。网站服务器可在http://www.cbs.dtu.dk/services/snpTree-1.0/上免费访问。
The advances and decreasing economical cost of whole genome sequencing (WGS), will soon make this technology available for routine infectious disease epidemiology. In epidemiological studies, outbreak isolates have very little diversity and require extensive genomic analysis to differentiate and classify isolates. One of the successfully and broadly used methods is analysis of single nucletide polymorphisms (SNPs). Currently, there are different tools and methods to identify SNPs including various options and cut-off values. Furthermore, all current methods require bioinformatic skills. Thus, we lack a standard and simple automatic tool to determine SNPs and construct phylogenetic tree from WGS data. Here we introduce snpTree, a server for online-automatic SNPs analysis. This tool is composed of different SNPs analysis suites, perl and python scripts. snpTree can identify SNPs and construct phylogenetic trees from WGS as well as from assembled genomes or contigs. WGS data in fastq format are aligned to reference genomes by BWA while contigs in fasta format are processed by Nucmer. SNPs are concatenated based on position on reference genome and a tree is constructed from concatenated SNPs using FastTree and a perl script. The online server was implemented by HTML, Java and python script. The server was evaluated using four published bacterial WGS data sets (V. cholerae, S. aureus CC398, S. Typhimurium and M. tuberculosis). The evalution results for the first three cases was consistent and concordant for both raw reads and assembled genomes. In the latter case the original publication involved extensive filtering of SNPs, which could not be repeated using snpTree. The snpTree server is an easy to use option for rapid standardised and automatic SNP analysis in epidemiological studies also for users with limited bioinformatic experience. The web server is freely accessible at http://www.cbs.dtu.dk/services/snpTree-1.0/.