Estimation of demo-genetic model probabilities with Approximate Bayesian Computation using linear discriminant analysis on summary statistics

Estimation of demo-genetic model probabilities with Approximate Bayesian Computation using linear discriminant analysis on summary statistics
复制标题

DOI:
10.1111/j.1755-0998.2012.03153.x
复制
发表时间:
2012-09-01
影响因子:
7.7
通讯作者:
Cornuet, Jean-Marie
Cornuet, Jean-Marie
中科院分区:
生物学1区
文献类型:
--
作者:
Estoup, Arnaud;Lombaert, Eric;Cornuet, Jean-Marie

文献摘要

被引文献

相似文献

使用近似贝叶斯计算(ABC)比较演示遗传模型是一个活跃的研究领域。尽管可以使用从各种标记类型获得的分子数据通过 ABC 来分析大量群体和模型(即场景),但当这些数字变得太大时,就会出现方法和计算问题。此外,罗伯特等人与等人类似。 (美国国家科学院院刊,2011, 108, 15112)表明,ABC 模型比较得出的结论本身不可信,需要额外的模拟分析。然而,当汇总统计量 (Ss) 和场景数量很大时,用于凭经验评估场景选择置信度的蒙特卡罗推理技术非常耗时。我们在这里描述了一种方法创新,在计算逻辑回归之前,使用 Ss 上的线性判别分析 (LDA) 来处理高效的 ABC 场景概率计算。与使用原始(即未经 LDA 转换的)Ss 的传统概率估计相比,我们使用模拟伪观测数据集 (pod) 来评估该方法的主要特征(精度和计算时间)。我们还在真实微卫星数据集上说明了该方法,以推断瓢虫异色瓢虫的入侵路线。我们发现从 LDA 转换的 Ss 和原始 Ss 计算出的场景概率具有很强的相关性。两种方法的 I 型和 II 型错误相似。我们观察到更快的概率计算(LDA 转换的 Ss 的速度增益约为 100 倍)大大提高了 ABC 从业者分析大量 Pod 的能力,因此提供了一种可管理的方法来凭经验评估可用于区分大量复杂场景的能力。
Comparison of demo-genetic models using Approximate Bayesian Computation (ABC) is an active research field. Although large numbers of populations and models (i.e. scenarios) can be analysed with ABC using molecular data obtained from various marker types, methodological and computational issues arise when these numbers become too large. Moreover, Robert et similar to al. (Proceedings of the National Academy of Sciences of the United States of America, 2011, 108, 15112) have shown that the conclusions drawn on ABC model comparison cannot be trusted per se and required additional simulation analyses. Monte Carlo inferential techniques to empirically evaluate confidence in scenario choice are very time-consuming, however, when the numbers of summary statistics (Ss) and scenarios are large. We here describe a methodological innovation to process efficient ABC scenario probability computation using linear discriminant analysis (LDA) on Ss before computing logistic regression. We used simulated pseudo-observed data sets (pods) to assess the main features of the method (precision and computation time) in comparison with traditional probability estimation using raw (i.e. not LDA transformed) Ss. We also illustrate the method on real microsatellite data sets produced to make inferences about the invasion routes of the coccinelid Harmonia axyridis. We found that scenario probabilities computed from LDA-transformed and raw Ss were strongly correlated. Type I and II errors were similar for both methods. The faster probability computation that we observed (speed gain around a factor of 100 for LDA-transformed Ss) substantially increases the ability of ABC practitioners to analyse large numbers of pods and hence provides a manageable way to empirically evaluate the power available to discriminate among a large set of complex scenarios.