Bayesian semi-supervised classification of bacterial samples using MLST databases.

Bayesian semi-supervised classification of bacterial samples using MLST databases.
复制标题

DOI:
10.1186/1471-2105-12-302
复制
发表时间:
2011-07-26
期刊:
影响因子:
3
通讯作者:
Corander J
Corander J
中科院分区:
生物学4区
文献类型:
--
作者:
Cheng L;Connor TR;Aanensen DM;Spratt BG;Corander J

文献摘要

参考文献

被引文献

相似文献

世界范围内的努力采样和表征的分子变异在大量的人类和动物病原体导致出现的多位点序列分型(MLST)数据库作为一个重要的工具,研究流行病学和病原体的演变。这些数据库中的许多数据库目前包含数千个多位点DNA序列类型(ST),其富含关于分离株的诸如血清型、抗生素抗性、宿主生物体等性状的元数据。因此,数据库的管理者有可能将病原体种群划分为代表不同进化谱系、地理相关组或其他亚群的子集,这些亚群根据数据库内的分子相似性和不相似性来定义。当与现有的元数据相结合时,这些子集可以为评估一组新的分离株相对于整个病原体群体的位置提供宝贵的信息。为了使MLST方案的用户能够用新的细菌分离株集查询数据库,并自动分析它们与现有策展序列的关系,我们在这里介绍了一种基于贝叶斯模型的MLST数据半监督分类方法。我们的方法可以使用MLST数据库作为训练集,并同时将任何一组查询序列分配到早期发现的谱系/种群中,同时还允许这些序列中的一些或全部形成先前未发现的遗传上不同的组。该工具提供了分类不确定性的概率量化,并且在计算上非常高效,从而能够快速分析大型数据库和查询序列集。后一个功能是通过MLST Web界面自动访问的必要先决条件。我们证明了我们的方法的多功能性,通过分析从MLST数据库的真实的和合成的数据。所介绍的用于查询ST集合的半监督分类的方法可在BAPS 5.4软件中免费获得用于Windows、Mac OS X和Linux操作系统,BAPS 5.4软件可在http://web.abo.fi/fak/mnf/mate/jc/software/baps.html下载。查询功能也可直接用于http://www.mlst.net上的金黄色葡萄球菌数据库,不久将可用于该门户网站上托管的其他菌种数据库。我们已经引入了一个基于模型的工具,用于新病原体样本的自动半监督分类,该工具可以集成到MLST数据库的Web界面中。特别地,当与现有元数据组合时,半监督标记可以提供用于评估新的查询菌株集合相对于由策展数据库表示的特定病原体群体的位置的宝贵信息。这些信息对于临床和基础研究都是有用的。
Worldwide effort on sampling and characterization of molecular variation within a large number of human and animal pathogens has lead to the emergence of multi-locus sequence typing (MLST) databases as an important tool for studying the epidemiology and evolution of pathogens. Many of these databases are currently harboring several thousands of multi-locus DNA sequence types (STs) enriched with metadata over traits such as serotype, antibiotic resistance, host organism etc of the isolates. Curators of the databases have thus the possibility of dividing the pathogen populations into subsets representing different evolutionary lineages, geographically associated groups, or other subpopulations, which are defined in terms of molecular similarities and dissimilarities residing within a database. When combined with the existing metadata, such subsets may provide invaluable information for assessing the position of a new set of isolates in relation to the whole pathogen population. To enable users of MLST schemes to query the databases with sets of new bacterial isolates and to automatically analyze their relation to existing curated sequences, we introduce here a Bayesian model-based method for semi-supervised classification of MLST data. Our method can use an MLST database as a training set and assign simultaneously any set of query sequences into the earlier discovered lineages/populations, while also allowing some or all of these sequences to form previously undiscovered genetically distinct groups. This tool provides probabilistic quantification of the classification uncertainty and is highly efficient computationally, thus enabling rapid analyses of large databases and sets of query sequences. The latter feature is a necessary prerequisite for an automated access through the MLST web interface. We demonstrate the versatility of our approach by anayzing both real and synthesized data from MLST databases. The introduced method for semi-supervised classification of sets of query STs is freely available for Windows, Mac OS X and Linux operative systems in BAPS 5.4 software which is downloadable at http://web.abo.fi/fak/mnf/mate/jc/software/baps.html. The query functionality is also directly available for the Staphylococcus aureus database at http://www.mlst.net and shortly will be available for other species databases hosted at this web portal. We have introduced a model-based tool for automated semi-supervised classification of new pathogen samples that can be integrated into the web interface of the MLST databases. In particular, when combined with the existing metadata, the semi-supervised labeling may provide invaluable information for assessing the position of a new set of query strains in relation to the particular pathogen population represented by the curated database. Such information will be useful both for clinical and basic research purposes.
DOI: 10.1128/jb.186.5.1518-1530.2004
发表时间: 2004-03-01
影响因子: 3.2
作者:
Feil, EJ;Li, BC;Spratt, BG
通讯作者: Spratt, BG
DOI: 10.1186/1471-2105-10-90
发表时间: 2009-03-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Marttinen P;Myllykangas S;Corander J
通讯作者: Corander J
DOI: 10.1186/1471-2156-11-94
发表时间: 2010-10-15
期刊: BMC genetics
影响因子: 2.9
作者:
Jombart T;Devillard S;Balloux F
通讯作者: Balloux F
DOI: 10.1186/1471-2105-5-86
发表时间: 2004-07-01
期刊: BMC bioinformatics
影响因子: 3
作者:
Jolley KA;Chan MS;Maiden MC
通讯作者: Maiden MC
DOI: 10.1371/journal.pcbi.1000455
发表时间: 2009-08
影响因子: 4.3
作者:
Tang J;Hanage WP;Fraser C;Corander J
通讯作者: Corander J