Flexible sampling of discrete data correlations without the marginal distributions

Flexible sampling of discrete data correlations without the marginal distributions
复制标题

DOI:
--
复制
发表时间:
2013-06
期刊:
--
影响因子:
--
通讯作者:
Freddie Kalaitzis;Ricardo Silva
Freddie Kalaitzis;Ricardo Silva
中科院分区:
其他
文献类型:
--
作者:
Freddie Kalaitzis;Ricardo Silva

文献摘要

相似文献

学习离散变量的联合依赖性是机器学习中的一个基本问题,其应用包括预测、聚类和降维。最近,Copula建模的框架由于其联合分布的模块化参数化而受到欢迎。在其他属性中,Copula提供了一个配方,用于将单变量边缘分布的灵活模型与适合于潜在高维依赖结构的参数族相结合。更根本的是,霍夫(2007)的扩展秩似然方法完全绕过了学习边际模型,当这些信息辅助手头的学习任务时,例如,标准降维问题或copula参数估计。其主要思想是通过可观察的秩统计来表示数据,忽略来自边缘的任何其他信息。推理通常是在贝叶斯框架中使用高斯Copula进行的,并且由于这意味着在约束数量随数据点数量二次增加的空间内进行采样而变得复杂。当使用现成的吉布斯采样时,结果是混合缓慢。我们提出了一个有效的算法的基础上,最新的进展约束哈密顿马尔可夫链蒙特卡罗,是简单的实现,不需要支付二次成本的样本量。
Learning the joint dependence of discrete variables is a fundamental problem in machine learning, with many applications including prediction, clustering and dimensionality reduction. More recently, the framework of copula modeling has gained popularity due to its modular parameterization of joint distributions. Among other properties, copulas provide a recipe for combining flexible models for univariate marginal distributions with parametric families suitable for potentially high dimensional dependence structures. More radically, the extended rank likelihood approach of Hoff (2007) bypasses learning marginal models completely when such information is ancillary to the learning task at hand as in, e.g., standard dimensionality reduction problems or copula parameter estimation. The main idea is to represent data by their observable rank statistics, ignoring any other information from the marginals. Inference is typically done in a Bayesian framework with Gaussian copulas, and it is complicated by the fact this implies sampling within a space where the number of constraints increases quadratically with the number of data points. The result is slow mixing when using off-the-shelf Gibbs sampling. We present an efficient algorithm based on recent advances on constrained Hamiltonian Markov chain Monte Carlo that is simple to implement and does not require paying for a quadratic cost in sample size.