Don’t Neglect the Obvious: On the Role of Unambiguous Words in Word Sense Disambiguation

Don’t Neglect the Obvious: On the Role of Unambiguous Words in Word Sense Disambiguation
复制标题

DOI:
10.18653/v1/2020.emnlp-main.283
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Daniel Loureiro;José Camacho-Collados
Daniel Loureiro;José Camacho-Collados
中科院分区:
其他
文献类型:
--
作者:
Daniel Loureiro;José Camacho-Collados

文献摘要

被引文献

相似文献

最先进的词义消歧(WSD)方法结合了两个不同的特征:预训练语言模型的强大功能和扩展此类模型覆盖范围的传播方法。这种传播是必要的,因为当前的语义注释语料库缺乏对底层语义库存(通常是 WordNet)中许多实例的覆盖。与此同时,明确的单词占 WordNet 中所有单词的很大一部分,而在现有的语义注释语料库中却很少被覆盖。在本文中,我们提出了一种简单的方法来为大型语料库中大多数明确的单词提供注释。我们介绍了 UWA(明确词注释)数据集,并展示了最先进的基于传播的模型如何使用它来大幅扩展其词义嵌入的覆盖范围和质量,从而改进其在 WSD 上的原始结果。
State-of-the-art methods for Word Sense Disambiguation (WSD) combine two different features: the power of pre-trained language models and a propagation method to extend the coverage of such models. This propagation is needed as current sense-annotated corpora lack coverage of many instances in the underlying sense inventory (usually WordNet). At the same time, unambiguous words make for a large portion of all words in WordNet, while being poorly covered in existing sense-annotated corpora. In this paper we propose a simple method to provide annotations for most unambiguous words in a large corpus. We introduce the UWA (Unambiguous Word Annotations) dataset and show how a state-of-the-art propagation-based model can use it to extend the coverage and quality of its word sense embeddings by a significant margin, improving on its original results on WSD.