Optimising predictive models to prioritise viral discovery in zoonotic reservoirs.

Optimising predictive models to prioritise viral discovery in zoonotic reservoirs.
复制标题

DOI:
10.1016/s2666-5247(21)00245-7
复制
发表时间:
2022-08
期刊:
The Lancet. Microbe
影响因子:
--
通讯作者:
Carlson CJ
Carlson CJ
中科院分区:
其他
文献类型:
--
作者:
Becker DJ;Albery GF;Sjodin AR;Poisot T;Bergner LM;Chen B;Cohen LE;Dallas TA;Eskew EA;Fagre AC;Farrell MJ;Guth S;Han BA;Simmons NB;Stock M;Teeling EC;Carlson CJ

文献摘要

被引文献

相似文献

尽管全球对“一个健康”疾病监测进行了投资,但识别和监测新型人畜共患病病毒的野生动物宿主仍然困难重重,成本高昂。统计模型可以指导采样目标的优先排序,但任何给定模型的预测可能是高度不确定的;此外,系统的模型验证是罕见的,因此,模型性能的驱动因素记录不足。在这里,我们使用β冠状病毒的蝙蝠宿主作为案例研究,用于比较和验证可能的储库宿主的预测模型的数据驱动过程。在2020年初,我们生成了一个由八个统计模型组成的集合,这些模型预测了宿主-病毒的关联,并为β冠状病毒的潜在蝙蝠宿主和SARS-CoV-2的桥梁宿主制定了优先采样建议。在一年多的时间内,我们跟踪发现了47种新的β冠状病毒蝙蝠宿主,验证了最初的预测,并动态更新了我们的分析流程。我们发现,基于生态特征的模型在预测这些新宿主方面表现良好,而网络方法在随机情况下的表现大致与预期一样好或更差。这些研究结果说明了集成建模的重要性,作为对混合模型质量的缓冲,并突出了预测模型中包括主机生态的价值。我们修改后的模型显示,与初始集合相比,性能有所改善,并预测全球有400多种蝙蝠物种可能是未检测到的β冠状病毒宿主。我们通过系统验证表明,机器学习模型可以帮助优化未发现病毒的野生动物采样,并说明如何通过预测,数据收集,验证和更新的动态过程来最好地实施这些方法。
Despite the global investment in One Health disease surveillance, it remains difficult and costly to identify and monitor the wildlife reservoirs of novel zoonotic viruses. Statistical models can guide sampling target prioritisation, but the predictions from any given model might be highly uncertain; moreover, systematic model validation is rare, and the drivers of model performance are consequently under-documented. Here, we use the bat hosts of betacoronaviruses as a case study for the data-driven process of comparing and validating predictive models of probable reservoir hosts. In early 2020, we generated an ensemble of eight statistical models that predicted host–virus associations and developed priority sampling recommendations for potential bat reservoirs of betacoronaviruses and bridge hosts for SARS-CoV-2. During a time frame of more than a year, we tracked the discovery of 47 new bat hosts of betacoronaviruses, validated the initial predictions, and dynamically updated our analytical pipeline. We found that ecological trait-based models performed well at predicting these novel hosts, whereas network methods consistently performed approximately as well or worse than expected at random. These findings illustrate the importance of ensemble modelling as a buffer against mixed-model quality and highlight the value of including host ecology in predictive models. Our revised models showed an improved performance compared with the initial ensemble, and predicted more than 400 bat species globally that could be undetected betacoronavirus hosts. We show, through systematic validation, that machine learning models can help to optimise wildlife sampling for undiscovered viruses and illustrates how such approaches are best implemented through a dynamic process of prediction, data collection, validation, and updating.