Mitigating Temporal-Drift: A Simple Approach to Keep NER Models Crisp

Mitigating Temporal-Drift: A Simple Approach to Keep NER Models Crisp
复制标题

DOI:
10.18653/v1/2021.socialnlp-1.14
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Shuguang Chen;Leonardo Neves;T. Solorio
Shuguang Chen;Leonardo Neves;T. Solorio
中科院分区:
其他
文献类型:
--
作者:
Shuguang Chen;Leonardo Neves;T. Solorio

文献摘要

相似文献

命名实体识别的神经模型的性能随着时间的推移而下降,变得陈旧。这种退化是由于时间漂移,即目标变量的统计属性随时间的变化。这个问题对于话题变化迅速的社交媒体数据来说尤其成问题。为了缓解这个问题,数据注释和模型的再训练是常见的。尽管它很有用,但这个过程是昂贵和耗时的,这激发了对有效模型更新的新研究。在本文中,我们提出了一种直观的方法来衡量推文的潜在趋势,并使用该度量来选择最具信息量的实例用于训练。我们在时态Twitter数据集上对三个最先进的模型进行了实验。我们的方法在训练数据较少的情况下,预测精度比其他方法有更大的提高,这使它成为一个有吸引力的、实用的解决方案。
Performance of neural models for named entity recognition degrades over time, becoming stale. This degradation is due to temporal drift, the change in our target variables’ statistical properties over time. This issue is especially problematic for social media data, where topics change rapidly. In order to mitigate the problem, data annotation and retraining of models is common. Despite its usefulness, this process is expensive and time-consuming, which motivates new research on efficient model updating. In this paper, we propose an intuitive approach to measure the potential trendiness of tweets and use this metric to select the most informative instances to use for training. We conduct experiments on three state-of-the-art models on the Temporal Twitter Dataset. Our approach shows larger increases in prediction accuracy with less training data than the alternatives, making it an attractive, practical solution.