Is Automated Topic Model Evaluation Broken?: The Incoherence of Coherence

Is Automated Topic Model Evaluation Broken?: The Incoherence of Coherence
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Alexander Miserlis Hoyle;Pranav Goel;Denis Peskov;Andrew Hian-Cheong;Jordan L. Boyd-Graber;P. Resnik
Alexander Miserlis Hoyle;Pranav Goel;Denis Peskov;Andrew Hian-Cheong;Jordan L. Boyd-Graber;P. Resnik
中科院分区:
其他
文献类型:
--
作者:
Alexander Miserlis Hoyle;Pranav Goel;Denis Peskov;Andrew Hian-Cheong;Jordan L. Boyd-Graber;P. Resnik

文献摘要

相似文献

与其他无监督方法的评估一样,主题模型评估可能会引起争议。然而,该领域已经围绕主题连贯性的自动估计进行了整合,这依赖于参考语料库中单词共现的频率。根据这些指标,当代神经主题模型超越了经典模型。与此同时,主题模型评估存在验证差距:为经典模型开发的自动一致性尚未通过神经模型的人体实验进行验证。此外,对主题建模文献的荟萃分析揭示了自动化主题建模基准中存在巨大的标准化差距。为了解决验证差距,我们将自动连贯性与两个最广泛接受的人类判断任务进行比较:主题评分和单词入侵。为了解决标准化差距,我们在两个常用数据集上系统地评估了一个占主导地位的经典模型和两个最先进的神经模型。当相应的人类评估没有宣布获胜模型时,自动评估宣布了获胜模型,这使人们对独立于人类判断的全自动评估的有效性提出了质疑。
Topic model evaluation, like evaluation of other unsupervised methods, can be contentious. However, the field has coalesced around automated estimates of topic coherence, which rely on the frequency of word co-occurrences in a reference corpus. Contemporary neural topic models surpass classical ones according to these metrics. At the same time, topic model evaluation suffers from a validation gap: automated coherence, developed for classical models, has not been validated using human experimentation for neural models. In addition, a meta-analysis of topic modeling literature reveals a substantial standardization gap in automated topic modeling benchmarks. To address the validation gap, we compare automated coherence with the two most widely accepted human judgment tasks: topic rating and word intrusion. To address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets. Automated evaluations declare a winning model when corresponding human evaluations do not, calling into question the validity of fully automatic evaluations independent of human judgments.