Making Heads and Tails of Models with Marginal Calibration for Sparse Tagsets
Making Heads and Tails of Models with Marginal Calibration for Sparse Tagsets
复制标题
DOI:
10.18653/v1/2021.findings-emnlp.423
复制
发表时间:
2021-09
期刊:
影响因子:
--
通讯作者:
Michael Kranzlein;Nelson F. Liu;Nathan Schneider
中科院分区:
文献类型:
--
作者:
Michael Kranzlein;Nelson F. Liu;Nathan Schneider
For interpreting the behavior of a probabilistic model, it is useful to measure a model's calibration--the extent to which it produces reliable confidence scores. We address the open problem of calibration for tagging models with sparse tagsets, and recommend strategies to measure and reduce calibration error (CE) in such models. We show that several post-hoc recalibration techniques all reduce calibration error across the marginal distribution for two existing sequence taggers. Moreover, we propose tag frequency grouping (TFG) as a way to measure calibration error in different frequency bands. Further, recalibrating each group separately promotes a more equitable reduction of calibration error across the tag frequency spectrum.