Objective priors from maximum entropy in data classification

Objective priors from maximum entropy in data classification
复制标题

DOI:
10.1016/j.inffus.2012.01.012
复制
发表时间:
2013-04
期刊:
Inf. Fusion
影响因子:
--
通讯作者:
F. Palmieri;D. Ciuonzo
F. Palmieri;D. Ciuonzo
中科院分区:
其他
文献类型:
--
作者:
F. Palmieri;D. Ciuonzo

文献摘要

被引文献

相似文献

在小数据集的分类问题中,缺乏先验分布的知识可能会使贝叶斯规则的应用受到质疑。统一或任意的先验可能提供分类答案,即使在简单的例子中,也可能最终与我们对问题的常识相矛盾。熵先验(EP),通过应用最大熵(ME)原则,似乎提供了很好的客观答案,在实际情况下,导致更保守的贝叶斯推理。EP的推导和适用于分类任务时,只有似然函数可用。在本文中,当推理仅基于一个样本时,我们还回顾了EP的使用,与从观测和类之间的互信息最大化获得的先验相比。这最后一个标准与KL分歧的最大化后验和先验之间的大样本集,导致众所周知的参考(或贝尔纳多)先验。我们对单个样本的比较考虑了这两种方法的前瞻性,并澄清了差异和潜力。EP的组合理由,启发沃利斯的熵定义的组合参数,也包括在内。EP的应用程序的序列(多个样本),可能会受到过度支配的类与最大熵也被认为是一个解决方案,保证后验一致性。一个明确的迭代算法提出EP确定完全从知识的似然函数。还包括模拟比较EP与均匀先验短序列。
Lack of knowledge of the prior distribution in classification problems that operate on small data sets may make the application of Bayes’ rule questionable. Uniform or arbitrary priors may provide classification answers that, even in simple examples, may end up contradicting our common sense about the problem. Entropic priors (EPs), via application of the maximum entropy (ME) principle, seem to provide good objective answers in practical cases leading to more conservative Bayesian inferences. EP are derived and applied to classification tasks when only the likelihood functions are available. In this paper, when inference is based only on one sample, we review the use of the EP also in comparison to priors that are obtained from maximization of the mutual information between observations and classes. This last criterion coincides with the maximization of the KL divergence between posteriors and priors that for large sample sets leads to the well-known reference (or Bernardo’s) priors. Our comparison on single samples considers both approaches in prospective and clarifies differences and potentials. A combinatorial justification for EP, inspired by Wallis’ combinatorial argument for entropy definition, is also included. The application of the EP to sequences (multiple samples) that may be affected by excessive domination of the class with the maximum entropy is also considered with a solution that guarantees posterior consistency. An explicit iterative algorithm is proposed for EP determination solely from knowledge of the likelihood functions. Simulations that compare EP with uniform priors on short sequences are also included.