AXIOMATIC DERIVATION OF THE PRINCIPLE OF MAXIMUM-ENTROPY AND THE PRINCIPLE OF MINIMUM CROSS-ENTROPY

AXIOMATIC DERIVATION OF THE PRINCIPLE OF MAXIMUM-ENTROPY AND THE PRINCIPLE OF MINIMUM CROSS-ENTROPY
复制标题

DOI:
10.1109/tit.1980.1056144
复制
发表时间:
1980-01-01
影响因子:
2.5
通讯作者:
JOHNSON, RW
JOHNSON, RW
中科院分区:
计算机科学2区
文献类型:
--
作者:
SHORE, JE;JOHNSON, RW

文献摘要

被引文献

相似文献

当新信息以期望值的形式给出时,Jaynes最大熵原理和Kullback最小交叉熵原理(最小有向散度)是唯一正确的归纳推理方法。以前的论证使用直观的论点,并依赖于熵和交叉熵的属性作为信息衡量标准。这里的方法假设,当有不同的方式考虑相同的信息时(例如,在不同的坐标系中),合理的归纳推理方法应该导致一致的结果。这一要求被形式化为四个一致性公理。这些都是以抽象信息运算符的形式陈述的,没有提到信息措施。证明了最大熵原理在如下意义上是正确的:最大化任何函数,除非该函数与熵具有相同的极大值,否则将导致不一致。换句话说,给定对期望值的约束形式的信息,只有一个满足约束的分布可以通过满足一致性公理的程序来选择;这种唯一的分布可以通过最大化熵来获得。这一结果既是直接建立的,也是作为最小交叉熵原理类似结果的特例(一致先验)建立的。得到了连续概率密度和离散分布的结果。
Jaynes's principle of maximum entropy and Kullbacks principle of minimum cross-entropy (minimum directed divergence) are shown to be uniquely correct methods for inductive inference when new information is given in the form of expected values. Previous justifications use intuitive arguments and rely on the properties of entropy and cross-entropy as information measures. The approach here assumes that reasonable methods of inductive inference should lead to consistent results when there are different ways of taking the same information into account (for example, in different coordinate system). This requirement is formalized as four consistency axioms. These are stated in terms of an abstract information operator and make no reference to information measures. It is proved that the principle of maximum entropy is correct in the following sense: maximizing any function but entropy will lead to inconsistency unless that function and entropy have identical maxima. In other words given information in the form of constraints on expected values, there is only one (distribution satisfying the constraints that can be chosen by a procedure that satisfies the consistency axioms; this unique distribution can be obtained by maximizing entropy. This result is established both directly and as a special case (uniform priors) of an analogous result for the principle of minimum cross-entropy. Results are obtained both for continuous probability densities and for discrete distributions.