Evaluating language models for the retrieval and categorization of lexical collocations

Evaluating language models for the retrieval and categorization of lexical collocations
复制标题

评估用于词汇搭配检索和分类的语言模型

DOI:
--
复制
发表时间:
2021
期刊:
Conference of the European Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
L. Wanner
L. Wanner
中科院分区:
--
文献类型:
--
作者:
Luis Espinosa Anke;Joan Codina;L. Wanner

文献摘要

被引文献

相似文献

词汇搭配是两个句法词汇的特质组合(例如,“大雨”或“迈出一步”)。应用程序。上下文中的搭配分为17个代表性的语义类别。尤其是如果搭配的第一个论点是主题,但通常无法区分同一语义类别中的不同句法结构,其次细粒的语义类别限制了给定碱基的一小部分有效相交的使用。
Lexical collocations are idiosyncratic combinations of two syntactically bound lexical items (e.g., “heavy rain” or “take a step”). Understanding their degree of compositionality and idiosyncrasy, as well their underlying semantics, is crucial for language learners, lexicographers and downstream NLP applications. In this paper, we perform an exhaustive analysis of current language models for collocation understanding. We first construct a dataset of apparitions of lexical collocations in context, categorized into 17 representative semantic categories. Then, we perform two experiments: (1) unsupervised collocate retrieval using BERT, and (2) supervised collocation classification in context. We find that most models perform well in distinguishing light verb constructions, especially if the collocation’s first argument acts as subject, but often fail to distinguish, first, different syntactic structures within the same semantic category, and second, fine-grained semantic categories which restrict the use of small sets of valid collocates for a given base.