Unbiased and efficient log-likelihood estimation with inverse binomial sampling.

Unbiased and efficient log-likelihood estimation with inverse binomial sampling.
复制标题

DOI:
10.1371/journal.pcbi.1008483
复制
发表时间:
2020-12
影响因子:
4.3
通讯作者:
Ma WJ
Ma WJ
中科院分区:
生物学2区
文献类型:
--
作者:
van Opheusden B;Acerbi L;Ma WJ

文献摘要

参考文献

被引文献

相似文献

科学假说的命运往往依赖于计算模型解释数据的能力,在现代统计方法中,这种能力通过似然函数来量化。对数似然是参数估计和模型评估的关键要素。然而,在计算生物学和神经科学等领域中,复杂模型的对数似然性通常难以通过分析或数值计算来计算。在这些情况下,研究人员通常只能通过将观察到的数据与模型模拟生成的合成观察值进行比较来估计对数可能性。通过模拟来近似可能性的标准技术要么使用数据的汇总统计量,要么有在估计中产生实质性偏差的风险。在这里,我们探索另一种方法,逆二项抽样(IBS),它可以有效地估计整个数据集的对数似然,并且没有偏差。对于每个观察,IBS从模拟器模型中抽取样本,直到一个与观察相匹配。对数似然估计是抽取样本数量的函数。这个估计量的方差是一致有界的,达到了无偏估计量的最小方差,并且我们可以计算方差的校准估计。我们提供了理论论据,有利于IBS和实证评估的方法,最大似然估计与基于模拟的模型。作为案例研究,我们从计算和认知神经科学中选取了三个日益复杂的模型拟合问题。在所有的问题中,IBS一般产生较低的误差,估计参数和最大对数似然值比其他抽样方法具有相同的平均样本数。我们的研究结果表明,IBS作为一个实用的,强大的,易于实现的方法,当精确的技术是不可用的对数似然评估的潜力。研究人员经常通过将数据与数学或计算模型的预测进行比较来验证科学假设。这种比较可以通过“对数似然”来量化,这是一个捕捉模型解释数据的程度的数字。然而,对于神经科学和计算生物学中常见的复杂模型,获得对数似然的精确公式可能是困难的。相反,对数似然通常通过模拟来自模型的合成观测(“采样”)来近似,并查看模拟数据与实际观测相匹配的频率。为了得出正确的科学结论,至关重要的是,这种抽样方法产生的对数似然估计是准确的(无偏的)。在这里,我们介绍了逆二项抽样(IBS),一种不同于传统方法的方法,从模型中提取的样本数量不是固定的,而是以一种简单的方式自适应调整。对于每个数据点,IBS从模型中采样,直到它与观察结果相匹配。我们表明,IBS是无偏的,并具有其他理想的统计特性,无论是理论上,并通过实证验证的三个案例研究,从计算和认知神经科学。在所有的例子中,IBS优于固定的抽样方法,证明IBS作为一个实用的,强大的,易于实现的对数似然评估方法的实用性。
The fate of scientific hypotheses often relies on the ability of a computational model to explain the data, quantified in modern statistical approaches by the likelihood function. The log-likelihood is the key element for parameter estimation and model evaluation. However, the log-likelihood of complex models in fields such as computational biology and neuroscience is often intractable to compute analytically or numerically. In those cases, researchers can often only estimate the log-likelihood by comparing observed data with synthetic observations generated by model simulations. Standard techniques to approximate the likelihood via simulation either use summary statistics of the data or are at risk of producing substantial biases in the estimate. Here, we explore another method, inverse binomial sampling (IBS), which can estimate the log-likelihood of an entire data set efficiently and without bias. For each observation, IBS draws samples from the simulator model until one matches the observation. The log-likelihood estimate is then a function of the number of samples drawn. The variance of this estimator is uniformly bounded, achieves the minimum variance for an unbiased estimator, and we can compute calibrated estimates of the variance. We provide theoretical arguments in favor of IBS and an empirical assessment of the method for maximum-likelihood estimation with simulation-based models. As case studies, we take three model-fitting problems of increasing complexity from computational and cognitive neuroscience. In all problems, IBS generally produces lower error in the estimated parameters and maximum log-likelihood values than alternative sampling methods with the same average number of samples. Our results demonstrate the potential of IBS as a practical, robust, and easy to implement method for log-likelihood evaluation when exact techniques are not available. Researchers often validate scientific hypotheses by comparing data with the predictions of a mathematical or computational model. This comparison can be quantified by the ‘log-likelihood’, a number that captures how well the model explains the data. However, for complex models common in neuroscience and computational biology, obtaining exact formulas for the log-likelihood can be difficult. Instead, the log-likelihood is usually approximated by simulating synthetic observations from the model (‘sampling’), and seeing how often the simulated data match the actual observations. To reach correct scientific conclusions, it is crucial that the log-likelihood estimates produced by such sampling methods are accurate (unbiased). Here, we introduce inverse binomial sampling (IBS), a method which differs from traditional approaches in that the number of samples drawn from the model is not fixed, but adaptively adjusted in a simple way. For each data point, IBS samples from the model until it matches the observation. We show that IBS is unbiased and has other desirable statistical properties, both theoretically and via empirical validation on three case studies from computational and cognitive neuroscience. Across all examples, IBS outperforms fixed sampling methods, demonstrating the utility of IBS as a practical, robust, and easy to implement method for log-likelihood evaluation.
DOI: 10.1162/106365603321828970
发表时间: 2003-03-01
影响因子: 6.8
作者:
Hansen, N;Muller, SD;Koumoutsakos, P
通讯作者: Koumoutsakos, P
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1214/aoms/1177706361
发表时间: 1959-01-01
影响因子: --
作者:
DEGROOT, MH
通讯作者: DEGROOT, MH
DOI: 10.2307/2332299
发表时间: 1945-11-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Haldane, JBS
通讯作者: Haldane, JBS
DOI: 10.1109/tevc.2008.924423
发表时间: 2009-02-01
影响因子: 14.3
作者:
Hansen, Nikolaus;Niederberger, Andre S. P.;Koumoutsakos, Petros
通讯作者: Koumoutsakos, Petros