Automatic Evaluation of Local Topic Quality

Automatic Evaluation of Local Topic Quality
复制标题

DOI:
10.18653/v1/p19-1076
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Jeffrey Lund;Piper Armstrong;Wilson Fearn;Stephen Cowley;Courtni Byun;Jordan L. Boyd-Graber;Kevin Seppi
Jeffrey Lund;Piper Armstrong;Wilson Fearn;Stephen Cowley;Courtni Byun;Jordan L. Boyd-Graber;Kevin Seppi
中科院分区:
其他
文献类型:
--
作者:
Jeffrey Lund;Piper Armstrong;Wilson Fearn;Stephen Cowley;Courtni Byun;Jordan L. Boyd-Graber;Kevin Seppi

文献摘要

相似文献

主题模型通常使用一致性等指标来评估它们生成的全局主题分布,但不考虑局部(令牌级)主题分配。令牌级分配对于分类等下游任务很重要。即使是最新的模型,旨在提高这些令牌级主题分配的质量,已经评估仅与全球指标。我们提出了一个任务,旨在引起人类的判断令牌级的主题分配。我们使用了各种主题模型类型和参数,发现全球指标与人类分配不一致。由于人类的评价是昂贵的,我们提出了各种自动化的指标来评估主题模型在本地一级。最后,我们将我们提出的指标与几个数据集上任务的人类判断相关联。我们发现,基于主题转换的百分比的评估与人类对本地主题质量的判断相关性最强。我们建议,这个新的指标,我们称之为一致性,通过与全球指标,如主题一致性评估新的主题模型时。
Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Even recent models, which aim to improve the quality of these token-level topic assignments, have been evaluated only with respect to global metrics. We propose a task designed to elicit human judgments of token-level topic assignments. We use a variety of topic model types and parameters and discover that global metrics agree poorly with human assignments. Since human evaluation is expensive we propose a variety of automated metrics to evaluate topic models at a local level. Finally, we correlate our proposed metrics with human judgments from the task on several datasets. We show that an evaluation based on the percent of topic switches correlates most strongly with human judgment of local topic quality. We suggest that this new metric, which we call consistency, be adopted alongside global metrics such as topic coherence when evaluating new topic models.