A new LDA formulation with covariates

A new LDA formulation with covariates
复制标题

一种新的带有协变量的 LDA 公式

DOI:
10.1080/03610918.2023.2171059
复制
发表时间:
2023
期刊:
Communications in Statistics - Simulation and Computation
影响因子:
--
通讯作者:
Valle, Denis
Valle, Denis
中科院分区:
--
文献类型:
--
作者:
Shimizu, Gilson Y.;Izbicki, Rafael;Valle, Denis

文献摘要

参考文献

相似文献

潜在狄利克雷位置(LDA)模型是一种用于创建混合成员聚类的流行方法。尽管LDA最初是为文本分析而开发的,但它已被广泛用于其他应用。我们提出了一个新的制定LDA模型,其中包括协变量。在这个模型中,LDA中嵌入了一个负二项回归,可以直接解释回归系数,并分析每个采样单元中特定于聚类的元素的数量(而不是像结构主题模型那样,分析重点是对每个聚类的比例进行建模)。我们使用切片抽样的吉布斯抽样算法来估计模型参数。我们依靠模拟来展示我们的算法如何能够成功地检索真实的参数值,以及使用协变量给出的信息来预测丰度矩阵的能力。该模型使用来自三个不同领域的真实的数据集进行说明:冠状病毒文章的文本挖掘,杂货店购物篮的分析,以及巴罗科罗拉多岛(巴拿马)的树种生态。该模型允许识别离散数据中的混合成员聚类,并提供协变量与这些聚类丰度之间关系的推断。
The Latent Dirichlet Location (LDA) model is a popular method for creating mixed-membership clusters. Despite having been originally developed for text analysis, LDA has been used for a wide range of other applications. We propose a new formulation for the LDA model which incorporates covariates. In this model, a negative binomial regression is embedded within LDA, enabling straight-forward interpretation of the regression coefficients and the analysis of the quantity of cluster-specific elements in each sampling units (instead of the analysis being focused on modeling the proportion of each cluster, as in Structural Topic Models). We use slice sampling within a Gibbs sampling algorithm to estimate model parameters. We rely on simulations to show how our algorithm is able to successfully retrieve the true parameter values and the ability to make predictions for the abundance matrix using the information given by the covariates. The model is illustrated using real data sets from three different areas: text-mining of Coronavirus articles, analysis of grocery shopping baskets, and ecology of tree species on Barro Colorado Island (Panama). This model allows the identification of mixed-membership clusters in discrete data and provides inference on the relationship between covariates and the abundance of these clusters.
DOI: --
发表时间: 2000-05
期刊: Genetics
影响因子: 3.3
作者:
J. K. Pritchard;Matthew Stephens;Peter Donnelly
通讯作者: J. K. Pritchard;Matthew Stephens;Peter Donnelly
改进医疗文档的主题表示以协助 COVID-19 文献探索
DOI: 10.18653/v1/2020.nlpcovid19-2.12
发表时间: 2020
期刊: Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020
影响因子: --
作者:
Yulia Otmakhova;K. Verspoor;Timothy Baldwin;Simon Suster;Jey Han Lau
通讯作者: Jey Han Lau
DOI: 10.1186/s12874-020-01059-y
发表时间: 2020-07-02
影响因子: 4
作者:
Liu, Nan;Chee, Marcel Lucas;Ong, Marcus Eng Hock
通讯作者: Ong, Marcus Eng Hock
DOI: 10.1016/j.knosys.2018.10.024
发表时间: 2019-01-01
影响因子: 8.8
作者:
Albuquerque, Pedro H. M.;do Valle, Denis Ribeiro;Li, Daijiang
通讯作者: Li, Daijiang
DOI: 10.1109/tpami.1984.4767596
发表时间: 1984-01-01
影响因子: 23.6
作者:
GEMAN, S;GEMAN, D
通讯作者: GEMAN, D