The effects of sampling strategy on the quality of reconstruction of viral population dynamics using Bayesian skyline family coalescent methods: A simulation study

The effects of sampling strategy on the quality of reconstruction of viral population dynamics using Bayesian skyline family coalescent methods: A simulation study
复制标题

DOI:
10.1093/ve/vew003
复制
发表时间:
2016-01-01
期刊:
影响因子:
5.3
通讯作者:
Rambaut, Andrew
Rambaut, Andrew
中科院分区:
医学2区
文献类型:
--
作者:
Hall, Matthew D.;Woolhouse, Mark E. J.;Rambaut, Andrew

文献摘要

被引文献

相似文献

病毒和其他病原体遗传数据总量的持续大规模增加导致了这样一种情况,即在系统发育分析中通常不可能包括每一个可用的序列,并期望该程序在合理的计算时间内完成。这就提出了如何选择一组序列进行分析的问题,特别是如果数据不仅仅用于推断系统发育树本身。分子流行病学抽样策略的设计一直是一个被忽视的研究领域。本文介绍了一个大规模的模拟演习,进行选择一个适当的策略时,使用GMRF skygrid,贝叶斯天际线家庭的聚结方法,以重建过去的人口动态。模拟情景旨在代表在多个地理位置对地方性病毒种群进行采样。在聚结或结构化聚结模型和从这些树模拟的序列下模拟大的聚结;然后根据各种方案对所得数据集进行下采样以进行分析。相同方案的不同重复之间的结果差异并非微不足道,因此,我们建议在可能的情况下,使用不同的数据集重复分析,以确定重建的元素不仅仅是所选特定样本集的结果。我们表明,一个单独的随机选择的序列可以引入虚假的行为,在中间线的Skygrid图,甚至边际似然估计可以表明复杂的动态,实际上并不存在。我们建议,中线不应被用来推断历史事件本身。与时间和空间位置(deme)的概率一致的抽样序列从来没有表现得比与当时和该位置的有效群体大小成比例的概率抽样更差,并且经常是上级。因此,我们建议在未来的研究设计中采用这种方法。我们还确认,在分析中包含来自单个地理位置的许多最近的序列往往会导致重建中的虚假瓶颈效应,并警告不要将其解释为真实的。
The ongoing large-scale increase in the total amount of genetic data for viruses and other pathogens has led to a situation in which it is often not possible to include every available sequence in a phylogenetic analysis and expect the procedure to complete in reasonable computational time. This raises questions about how a set of sequences should be selected for analysis, particularly if the data are used to infer more than just the phylogenetic tree itself. The design of sampling strategies for molecular epidemiology has been a neglected field of research. This article describes a large-scale simulation exercise that was undertaken to select an appropriate strategy when using the GMRF skygrid, one of the Bayesian skyline family of coalescent methods, in order to reconstruct past population dynamics. The simulated scenarios were intended to represent sampling for the population of an endemic virus across multiple geographical locations. Large phylogenies were simulated under a coalescent or structured coalescent model and sequences simulated from these trees; the resulting datasets were then downsampled for analyses according to a variety of schemes. Variation in results between different replicates of the same scheme was not insignificant, and as a result, we recommend that where possible analyses are repeated with different datasets in order to establish that elements of a reconstruction are not simply the result of the particular set of samples selected. We show that an individual stochastic choice of sequences can introduce spurious behaviour in the median line of the skygrid plot and that even marginal likelihood estimation can suggest complicated dynamics that were not in fact present. We recommend that the median line should not be used to infer historical events on its own. Sampling sequences with uniform probability with respect to both time and spatial location (deme) never performed worse than sampling with probability proportional to the effective population size at that time and in that location and frequently was superior. As a result, we recommend this approach in the design of future studies. We also confirm that the inclusion of many recent sequences from a single geographical location in an analysis tends to result in a spurious bottleneck effect in the reconstruction and caution against interpreting this as genuine.