Double Spike Dirichlet Priors for Structured Weighting

Double Spike Dirichlet Priors for Structured Weighting
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Huiming Lin;Meng Li
Huiming Lin;Meng Li
中科院分区:
其他
文献类型:
--
作者:
Huiming Lin;Meng Li

文献摘要

相似文献

在各种各样的应用程序中,为大量对象分配权重是一项基本任务。在本文中,我们引入了结构化高维概率简单体的概念,它的大部分成分是零或接近零,其余成分彼此接近。这种结构很好地受到以下因素的驱动:1)现代应用中常见的高维权重,以及2)无处不在的例子,在这些例子中,相等的权重——尽管很简单——经常获得有利的甚至是最先进的预测性能。然而,这种特殊的结构在计算和统计上都提出了独特的挑战。为了解决这些挑战,我们提出了一类新的双尖峰狄利克雷先验,将概率单纯形缩小为具有所需结构的概率单纯形。当应用于集成学习时,这种先验导致了结构化高维集成的贝叶斯方法,这对于预测组合和改进随机森林非常有用,同时实现了不确定性量化。我们设计了高效的马尔可夫链蒙特卡罗算法,便于实现。建立后宫收缩率以提供理论支持。我们通过模拟和使用欧洲中央银行专业预测者调查数据集和UCI数据集的两个真实数据应用,证明了所提出方法的广泛适用性和竞争力。
Assigning weights to a large pool of objects is a fundamental task in a wide variety of applications. In this article, we introduce a concept of structured high-dimensional probability simplexes, whose most components are zero or near zero and the remaining ones are close to each other. Such structure is well motivated by 1) high-dimensional weights that are common in modern applications, and 2) ubiquitous examples in which equal weights---despite their simplicity---often achieve favorable or even state-of-the-art predictive performances. This particular structure, however, presents unique challenges both computationally and statistically. To address these challenges, we propose a new class of double spike Dirichlet priors to shrink a probability simplex to one with the desired structure. When applied to ensemble learning, such priors lead to a Bayesian method for structured high-dimensional ensembles that is useful for forecast combination and improving random forests, while enabling uncertainty quantification. We design efficient Markov chain Monte Carlo algorithms for easy implementation. Posterior contraction rates are established to provide theoretical support. We demonstrate the wide applicability and competitive performance of the proposed methods through simulations and two real data applications using the European Central Bank Survey of Professional Forecasters dataset and a UCI dataset.