Making Heads and Tails of Models with Marginal Calibration for Sparse Tagsets

Making Heads and Tails of Models with Marginal Calibration for Sparse Tagsets
复制标题

DOI:
10.18653/v1/2021.findings-emnlp.423
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
Michael Kranzlein;Nelson F. Liu;Nathan Schneider
Michael Kranzlein;Nelson F. Liu;Nathan Schneider
中科院分区:
其他
文献类型:
--
作者:
Michael Kranzlein;Nelson F. Liu;Nathan Schneider

文献摘要

相似文献

为了解释概率模型的行为,测量模型的校准(模型产生可靠置信度分数的程度)很有用。我们解决了具有稀疏标签集的标签模型的校准开放问题,并推荐了测量和减少此类模型中的校准误差(CE)的策略。我们表明,几种事后重新校准技术都可以减少两个现有序列标记器的边缘分布的校准误差。此外,我们提出标签频率分组(TFG)作为测量不同频段校准误差的方法。此外,单独重新校准每个组可以更公平地减少标签频谱上的校准误差。
For interpreting the behavior of a probabilistic model, it is useful to measure a model's calibration--the extent to which it produces reliable confidence scores. We address the open problem of calibration for tagging models with sparse tagsets, and recommend strategies to measure and reduce calibration error (CE) in such models. We show that several post-hoc recalibration techniques all reduce calibration error across the marginal distribution for two existing sequence taggers. Moreover, we propose tag frequency grouping (TFG) as a way to measure calibration error in different frequency bands. Further, recalibrating each group separately promotes a more equitable reduction of calibration error across the tag frequency spectrum.