Evaluating language models for the retrieval and categorization of lexical collocations
Evaluating language models for the retrieval and categorization of lexical collocations
复制标题
评估用于词汇搭配检索和分类的语言模型
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
L. Wanner
中科院分区:
文献类型:
--
作者:
Luis Espinosa Anke;Joan Codina;L. Wanner
Lexical collocations are idiosyncratic combinations of two syntactically bound lexical items (e.g., “heavy rain” or “take a step”). Understanding their degree of compositionality and idiosyncrasy, as well their underlying semantics, is crucial for language learners, lexicographers and downstream NLP applications. In this paper, we perform an exhaustive analysis of current language models for collocation understanding. We first construct a dataset of apparitions of lexical collocations in context, categorized into 17 representative semantic categories. Then, we perform two experiments: (1) unsupervised collocate retrieval using BERT, and (2) supervised collocation classification in context. We find that most models perform well in distinguishing light verb constructions, especially if the collocation’s first argument acts as subject, but often fail to distinguish, first, different syntactic structures within the same semantic category, and second, fine-grained semantic categories which restrict the use of small sets of valid collocates for a given base.