Dependent Multinomial Models Made Easy: Stick-Breaking with the Polya-gamma Augmentation

Dependent Multinomial Models Made Easy: Stick-Breaking with the Polya-gamma Augmentation
复制标题

DOI:
--
复制
发表时间:
2015-06
期刊:
--
影响因子:
--
通讯作者:
Scott W. Linderman;Matthew J. Johnson;Ryan P. Adams
Scott W. Linderman;Matthew J. Johnson;Ryan P. Adams
中科院分区:
其他
文献类型:
--
作者:
Scott W. Linderman;Matthew J. Johnson;Ryan P. Adams

文献摘要

被引文献

相似文献

许多实际建模问题涉及离散数据,这些数据最好表示为从多项或分类分布中抽取的数据。例如,DNA 序列中的核苷酸、给定州和年份的儿童姓名以及文本文档通常都使用多项分布进行建模。在所有这些情况下,我们预计绘图之间存在某种形式的依赖性:DNA 链中一个位置的核苷酸可能依赖于前面的核苷酸,儿童的名字每年都高度相关,文本中的主题可能是相关的和动态的。典型的狄利克雷多项式公式并不能自然地捕获这些依赖性。在这里,我们利用逻辑断棒表示和 Polya-gamma 增强的最新创新,根据潜在变量和联合高斯似然重新表述多项分布,使我们能够以最小的开销利用大量贝叶斯推理技术来构建高斯模型。
Many practical modeling problems involve discrete data that are best represented as draws from multinomial or categorical distributions. For example, nucleotides in a DNA sequence, children's names in a given state and year, and text documents are all commonly modeled with multinomial distributions. In all of these cases, we expect some form of dependency between the draws: the nucleotide at one position in the DNA strand may depend on the preceding nucleotides, children's names are highly correlated from year to year, and topics in text may be correlated and dynamic. These dependencies are not naturally captured by the typical Dirichlet-multinomial formulation. Here, we leverage a logistic stick-breaking representation and recent innovations in Polya-gamma augmentation to reformulate the multinomial distribution in terms of latent variables with jointly Gaussian likelihoods, enabling us to take advantage of a host of Bayesian inference techniques for Gaussian models with minimal overhead.