Likelihood estimation of sparse topic distributions in topic models and its applications to Wasserstein document distance calculations

Likelihood estimation of sparse topic distributions in topic models and its applications to Wasserstein document distance calculations
复制标题

DOI:
10.1214/22-aos2229
复制
发表时间:
2021-07
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
Xin Bing;F. Bunea;Seth Strimas-Mackey;M. Wegkamp
Xin Bing;F. Bunea;Seth Strimas-Mackey;M. Wegkamp
中科院分区:
其他
文献类型:
--
作者:
Xin Bing;F. Bunea;Seth Strimas-Mackey;M. Wegkamp

文献摘要

被引文献

相似文献

本文研究主题模型中高维、离散、可能稀疏的混合模型的估计问题。数据由在$n$个独立文档中观察到的$p$个单词的多项计数组成。在主题模型中,假设$p\ * n$期望词频矩阵被分解为$p\ * K$词-主题矩阵$ a $和$K\ * n$主题-文档矩阵$T$。由于这两个矩阵的列都表示属于简单概率的条件概率,因此$A$的列被视为所有文档共有的$p$维混合分量,而$T$的列被视为特定于文档的$K$维混合权重,并且允许是稀疏的。主要的兴趣是当$A$已知或未知时,为混合权$T$的估计量提供尖锐的,有限样本的$\ell_1$范数收敛率。对于已知的$A$,我们建议对$T$进行最大似然估计。我们对MLE的非标准分析不仅确定了它的收敛速度,而且揭示了一个显著的性质:不需要额外的正则化,MLE可以是精确稀疏的,并且包含T的真零模式。我们进一步证明了在一大类稀疏主题分布中,MLE既具有极大极小最优性,又具有对未知稀疏度的自适应性。当$A$未知时,我们通过优化对应于$A$的插入式通用估计器$\hat{A}$的似然函数来估计$T$。对于任何估计量$\hat{A}$满足接近$A$的详细条件,结果$T$的估计量被证明保留了为MLE建立的性质。允许环境尺寸$K$和$p$随样本量的增大而增大。我们的应用是估计文档生成分布之间的1-Wasserstein距离。我们提出,估计和分析两个概率文档表示之间的新1-Wasserstein距离。
This paper studies the estimation of high-dimensional, discrete, possibly sparse, mixture models in topic models. The data consists of observed multinomial counts of $p$ words across $n$ independent documents. In topic models, the $p\times n$ expected word frequency matrix is assumed to be factorized as a $p\times K$ word-topic matrix $A$ and a $K\times n$ topic-document matrix $T$. Since columns of both matrices represent conditional probabilities belonging to probability simplices, columns of $A$ are viewed as $p$-dimensional mixture components that are common to all documents while columns of $T$ are viewed as the $K$-dimensional mixture weights that are document specific and are allowed to be sparse. The main interest is to provide sharp, finite sample, $\ell_1$-norm convergence rates for estimators of the mixture weights $T$ when $A$ is either known or unknown. For known $A$, we suggest MLE estimation of $T$. Our non-standard analysis of the MLE not only establishes its $\ell_1$ convergence rate, but reveals a remarkable property: the MLE, with no extra regularization, can be exactly sparse and contain the true zero pattern of $T$. We further show that the MLE is both minimax optimal and adaptive to the unknown sparsity in a large class of sparse topic distributions. When $A$ is unknown, we estimate $T$ by optimizing the likelihood function corresponding to a plug in, generic, estimator $\hat{A}$ of $A$. For any estimator $\hat{A}$ that satisfies carefully detailed conditions for proximity to $A$, the resulting estimator of $T$ is shown to retain the properties established for the MLE. The ambient dimensions $K$ and $p$ are allowed to grow with the sample sizes. Our application is to the estimation of 1-Wasserstein distances between document generating distributions. We propose, estimate and analyze new 1-Wasserstein distances between two probabilistic document representations.