Machine learning to predict the source of campylobacteriosis using whole genome data.
Machine learning to predict the source of campylobacteriosis using whole genome data.
复制标题
DOI:
10.1371/journal.pgen.1009436
复制
发表时间:
2021-10
期刊:
影响因子:
4.5
通讯作者:
Wilson DJ
中科院分区:
文献类型:
--
作者:
Arning N;Sheppard SK;Bayliss S;Clifton DA;Wilson DJ
Campylobacteriosis is among the world’s most common foodborne illnesses, caused predominantly by the bacterium Campylobacter jejuni. Effective interventions require determination of the infection source which is challenging as transmission occurs via multiple sources such as contaminated meat, poultry, and drinking water. Strain variation has allowed source tracking based upon allelic variation in multi-locus sequence typing (MLST) genes allowing isolates from infected individuals to be attributed to specific animal or environmental reservoirs. However, the accuracy of probabilistic attribution models has been limited by the ability to differentiate isolates based upon just 7 MLST genes. Here, we broaden the input data spectrum to include core genome MLST (cgMLST) and whole genome sequences (WGS), and implement multiple machine learning algorithms, allowing more accurate source attribution. We increase attribution accuracy from 64% using the standard iSource population genetic approach to 71% for MLST, 85% for cgMLST and 78% for kmerized WGS data using the classifier we named aiSource. To gain insight beyond the source model prediction, we use Bayesian inference to analyse the relative affinity of C. jejuni strains to infect humans and identified potential differences, in source-human transmission ability among clonally related isolates in the most common disease causing lineage (ST-21 clonal complex). Providing generalizable computationally efficient methods, based upon machine learning and population genetics, we provide a scalable approach to global disease surveillance that can continuously incorporate novel samples for source attribution and identify fine-scale variation in transmission potential. C. jejuni are the most common cause of food-borne bacterial gastroenteritis but the relative contribution of different sources is incompletely understood. We traced the origin of human C. jejuni infections using machine learning algorithms that compare the DNA sequences of bacteria sampled from infected people, contaminated chickens, cattle, sheep, wild birds, and the environment. This approach achieved improvement in accuracy of source attribution by 33% over existing methods that use only a subset of genes within the genome and provided evidence for the relative contribution of different infection sources. Sometimes even very similar bacteria showed differences, demonstrating the value of basing analyses on the entire genome when developing this algorithm that can be used for understanding the global epidemiology and other important bacterial infections.
登录
查看更多内容
影响因子:
--
作者:
Jolley KA;Bray JE;Maiden MCJ
通讯作者:
Maiden MCJ
影响因子:
4.4
作者:
Chen, Xi;Ishwaran, Hemant
通讯作者:
Ishwaran, Hemant
DOI:
10.1038/ismej.2015.149
发表时间:
2016-03
期刊:
The ISME journal
影响因子:
--
作者:
Dearlove BL;Cody AJ;Pascoe B;Méric G;Wilson DJ;Sheppard SK
通讯作者:
Sheppard SK
影响因子:
3.3
作者:
Ansari MA;Didelot X
通讯作者:
Didelot X
影响因子:
4.6
作者:
Kirk, Karina Frahm;Meric, Guillaume;Nielsen, Henrik
通讯作者:
Nielsen, Henrik