The advantage of intergenic regions as genomic features for machine-learning-based host attribution of Salmonella Typhimurium from the USA.

The advantage of intergenic regions as genomic features for machine-learning-based host attribution of Salmonella Typhimurium from the USA.
复制标题

DOI:
10.1099/mgen.0.001116
复制
发表时间:
2023-10
期刊:
影响因子:
3.9
通讯作者:
Gally, David L.
Gally, David L.
中科院分区:
生物学2区
文献类型:
--
作者:
Chalka, Antonia;Dallman, Tim J.;Vohra, Prerna;Stevens, Mark P.;Gally, David L.

文献摘要

参考文献

相似文献

肠道沙门氏菌 是一种分类学上多样化的病原体,具有超过2600个血清型,与包括人类、其他哺乳动物、鸟类和爬行动物在内的多种动物宿主相关。一些血清型是宿主特异性的或宿主限制性的,并在不同的宿主物种中引起疾病,而另一些血清型,如血清型S。鼠伤寒沙门氏菌(STm)是多面手,有可能定殖各种各样的物种。然而,即使在多面手血清型如STm中,也越来越清楚存在嗜性和毒力不同的致病变体。识别宿主特异性背后的遗传因素是复杂的,但数千个基因组序列的可用性和机器学习的进步使建立特定宿主预测模型成为可能,以帮助控制疫情,并预测来自动物和其他宿主的分离株对人类的致病潜力。我们通过构建在广泛的基因组特征上训练的宿主关联预测模型,并将其与基于最近邻同源性的预测进行比较,推进了这一领域。SNPs,蛋白质变体(PV),抗菌素耐药性(AMR)谱和基因间区域(IGR)从3883个高质量的STm组件中提取,这些组件收集自美国的人类,猪,牛和家禽,并用于构建随机森林(RF)机器学习模型。来自农场动物的另外244个最近的STm组件用作进一步确认的测试集。基于PV和IGR的模型在预测分离株的起源宿主方面具有最佳性能,并且优于最近邻系统发育宿主预测以及基于SNP或AMR数据的模型。然而,当用与训练集在遗传学上不同的分离株进行测试时,模型没有产生可靠的预测。IGR和PV模型通常能够区分集群中的人分离株,其中大多数分离株来自单一动物来源。值得注意的是,IGR是在多种模型中具有最佳性能的特征,这可能是由于IGR既作为其侧翼基因的代表,相当于PV,同时也捕获基因组调控变异,如改变的启动子区。IGR和PV模型预测,美国约45%的STm人类感染来自牛,约40%来自家禽,约14.5%来自猪,尽管来自其他来源的分离株序列未用于训练。   总之,该研究表明,与基于SNP和核心基因组的预测相比,当应用于现有的群体结构时,以IGR和PV作为特征的模型的准确性显著提高。本文包含由Microreact托管的数据。
Salmonella enterica is a taxonomically diverse pathogen with over 2600 serovars associated with a wide variety of animal hosts including humans, other mammals, birds and reptiles. Some serovars are host-specific or host-restricted and cause disease in distinct host species, while others, such as serovar S. Typhimurium (STm), are generalists and have the potential to colonize a wide variety of species. However, even within generalist serovars such as STm it is becoming clear that pathovariants exist that differ in tropism and virulence. Identifying the genetic factors underlying host specificity is complex, but the availability of thousands of genome sequences and advances in machine learning have made it possible to build specific host prediction models to aid outbreak control and predict the human pathogenic potential of isolates from animals and other reservoirs. We have advanced this area by building host-association prediction models trained on a wide range of genomic features and compared them with predictions based on nearest-neighbour phylogeny. SNPs, protein variants (PVs), antimicrobial resistance (AMR) profiles and intergenic regions (IGRs) were extracted from 3883 high-quality STm assemblies collected from humans, swine, bovine and poultry in the USA, and used to construct Random Forest (RF) machine learning models. An additional 244 recent STm assemblies from farm animals were used as a test set for further validation. The models based on PVs and IGRs had the best performance in terms of predicting the host of origin of isolates and outperformed nearest-neighbour phylogenetic host prediction as well as models based on SNPs or AMR data. However, the models did not yield reliable predictions when tested with isolates that were phylogenetically distinct from the training set. The IGR and PV models were often able to differentiate human isolates in clusters where the majority of isolates were from a single animal source. Notably, IGRs were the feature with the best performance across multiple models which may be due to IGRs acting as both a representation of their flanking genes, equivalent to PVs, while also capturing genomic regulatory variation, such as altered promoter regions. The IGR and PV models predict that ~45 % of the human infections with STm in the USA originate from bovine, ~40 % from poultry and ~14.5 % from swine, although sequences of isolates from other sources were not used for training. In summary, the research demonstrates a significant gain in accuracy for models with IGRs and PVs as features compared to SNP-based and core genome phylogeny predictions when applied within the existing population structure. This article contains data hosted by Microreact.
DOI: 10.1371/journal.pmed.1001923
发表时间: 2015-12
期刊: PLoS medicine
影响因子: 15.8
作者:
Havelaar AH;Kirk MD;Torgerson PR;Gibb HJ;Hald T;Lake RJ;Praet N;Bellinger DC;de Silva NR;Gargouri N;Speybroeck N;Cawthorne A;Mathers C;Stein C;Angulo FJ;Devleesschauwer B;World Health Organization Foodborne Disease Burden Epidemiology Reference Group
通讯作者: World Health Organization Foodborne Disease Burden Epidemiology Reference Group
DOI: 10.1016/j.resmic.2014.07.004
发表时间: 2014-09-01
影响因子: 2.6
作者:
Issenhuth-Jeanjean, Sylvie;Roggentin, Peter;Weill, Francois-Xavier
通讯作者: Weill, Francois-Xavier
DOI: 10.1371/journal.pone.0046675
发表时间: 2012-12-21
期刊: PLOS ONE
影响因子: 3.7
作者:
Hennebry, Sarah C.;Sait, Leanne C.;Strugnell, Richard A.
通讯作者: Strugnell, Richard A.
DOI: 10.1186/s13059-016-1108-8
发表时间: 2016-11-25
期刊: Genome biology
影响因子: 12.3
作者:
Brynildsrud O;Bohlin J;Scheffer L;Eldholm V
通讯作者: Eldholm V
DOI: 10.1371/journal.pone.0177459
发表时间: 2017
期刊: PloS one
影响因子: 3.7
作者:
Kurtzer GM;Sochat V;Bauer MW
通讯作者: Bauer MW