Distributed multinomial regression

Distributed multinomial regression
复制标题

DOI:
10.1214/15-aoas831
复制
发表时间:
2013-11
期刊:
The Annals of Applied Statistics
影响因子:
--
通讯作者:
Matt Taddy
Matt Taddy
中科院分区:
其他
文献类型:
--
作者:
Matt Taddy

文献摘要

被引文献

相似文献

本文介绍了一种基于模型的方法来分布式计算多项逻辑回归(softmax)。我们将每个响应类别的计数视为独立的泊松回归,通过插件估计跨类别共享的固定效应。这项工作是由用于分析大量随机计数的高维响应多项式模型驱动的。我们的激励应用程序是在文本分析中,文档标记化和令牌计数建模为从依赖于文档属性的多项式产生。我们估计这样的模型,从Yelp的评论的公开可用的数据集,与文本回归到一个大的解释变量集(用户,业务和评级信息)。拟合的模型作为探索单词和感兴趣的变量之间的联系,减少维度到监督因子得分,并进行预测的基础。我们认为,这里的方法提供了一个有吸引力的选择,社会科学家和其他文本分析师谁希望把熟悉的回归工具,以承担文本数据。
This article introduces a model-based approach to distributed computing for multinomial logistic (softmax) regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that are used in analysis of a large number of random counts. Our motivating applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available data set of reviews from Yelp, with text regressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of interest, for reducing dimension into supervised factor scores, and for prediction. We argue that the approach herein provides an attractive option for social scientists and other text analysts who wish to bring familiar regression tools to bear on text data.