FUBAR: A Fast, Unconstrained Bayesian AppRoximation for Inferring Selection

FUBAR: A Fast, Unconstrained Bayesian AppRoximation for Inferring Selection
复制标题

DOI:
10.1093/molbev/mst030
复制
发表时间:
2013-05-01
影响因子:
10.7
通讯作者:
Scheffler, Konrad
Scheffler, Konrad
中科院分区:
生物学1区
文献类型:
--
作者:
Murrell, Ben;Moola, Sasha;Scheffler, Konrad

文献摘要

被引文献

相似文献

基于模型的自然选择分析通常将站点分类为相对较少的站点类。强制每个站点属于这些类别之一,对选择参数的分布施加了不切实际的约束,这可能会导致由于模型错误指定而导致的误导性推断。我们提出了一个近似的分层贝叶斯方法,使用马尔可夫链蒙特卡罗(MCMC)例程,确保对模型的错误指定的鲁棒性平均超过大量的预定义的网站类。这使得选择参数的分布基本上不受约束,并且还允许经历正选择和纯化选择的位点比现有方法更快地被确定数量级。我们表明,流行的随机效应的可能性方法可以产生误导性的结果时,分配到同一网站类的网站经历不同程度的积极或净化选择-一个不可避免的情况下,使用少量的网站类。我们的快速无约束贝叶斯估计(FUBAR)不受此问题的影响,同时实现比现有的无约束(固定效应似然)方法更高的功率。FUBAR的速度优势使我们能够分析比其他方法更大的数据集:我们在大型流感血凝素数据集(3,142个序列)上说明了这一点。FUBAR可以作为批处理文件在最新的HyPhy发行版(http://www.example.com)以及Datamonkey Web服务器(http://www.datamonkey.org/)上获得。www.hyphy.org
Model-based analyses of natural selection often categorize sites into a relatively small number of site classes. Forcing each site to belong to one of these classes places unrealistic constraints on the distribution of selection parameters, which can result in misleading inference due to model misspecification. We present an approximate hierarchical Bayesian method using a Markov chain Monte Carlo (MCMC) routine that ensures robustness against model misspecification by averaging over a large number of predefined site classes. This leaves the distribution of selection parameters essentially unconstrained, and also allows sites experiencing positive and purifying selection to be identified orders of magnitude faster than by existing methods. We demonstrate that popular random effects likelihood methods can produce misleading results when sites assigned to the same site class experience different levels of positive or purifying selection-an unavoidable scenario when using a small number of site classes. Our Fast Unconstrained Bayesian AppRoximation (FUBAR) is unaffected by this problem, while achieving higher power than existing unconstrained (fixed effects likelihood) methods. The speed advantage of FUBAR allows us to analyze larger data sets than other methods: We illustrate this on a large influenza hemagglutinin data set (3,142 sequences). FUBAR is available as a batch file within the latest HyPhy distribution (http://www.hyphy.org), as well as on the Datamonkey web server ( http://www.datamonkey.org/).