Low-probability states, data statistics, and entropy estimation

Low-probability states, data statistics, and entropy estimation
复制标题

低概率状态、数据统计和熵估计

DOI:
10.1103/physreve.108.014101
复制
发表时间:
2023
期刊:
影响因子:
2.4
通讯作者:
Nemenman, Ilya
Nemenman, Ilya
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Hernández, Damián G.;Roman, Ahmed;Nemenman, Ilya

文献摘要

相似文献

复杂系统分析中的一个基本问题是对系统在状态空间上的概率分布的熵进行可靠的估计。这是困难的,因为未采样的状态可以对熵有很大的贡献,而它们对熵的最大似然估计没有贡献,熵的最大似然估计用观察到的频率代替概率。贝叶斯估计通过引入概率分布的低概率尾部模型克服了这一障碍。观测数据的哪些统计特征决定了尾部的模型,从而决定了这些估计量的输出,目前还不清楚。在这里,我们表明,著名的离散状态空间上的概率分布的熵估计模型的低概率尾部的结构主要基于几个统计数据:样本大小,最大似然估计,样本之间的巧合的数量,和分散的巧合。基于这些统计数据,我们推导出欠采样分布的近似解析熵估计,并利用这些结果对贝叶斯熵估计的工作原理提出了直观的理解。
A fundamental problem in the analysis of complex systems is getting a reliable estimate of the entropy of their probability distributions over the state space. This is difficult because unsampled states can contribute substantially to the entropy, while they do not contribute to the maximum likelihood estimator of entropy, which replaces probabilities by the observed frequencies. Bayesian estimators overcome this obstacle by introducing a model of the low-probability tail of the probability distribution. Which statistical features of the observed data determine the model of the tail, and hence the output of such estimators, remains unclear. Here we show that well-known entropy estimators for probability distributions on discrete state spaces model the structure of the low-probability tail based largely on a few statistics of the data: the sample size, the maximum likelihood estimate, the number of coincidences among the samples, and the dispersion of the coincidences. We derive approximate analytical entropy estimators for undersampled distributions based on these statistics, and we use the results to propose an intuitive understanding of how the Bayesian entropy estimators work.