SemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding

SemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding
复制标题

SemEval-2022 任务 2:多语言惯用语检测和句子嵌入

DOI:
10.48550/arxiv.2204.10050
复制
发表时间:
2022
影响因子:
16.6
通讯作者:
Aline Villavicencio
Aline Villavicencio
中科院分区:
计算机科学1区
文献类型:
--
作者:
Harish Tayyar Madabushi;Edward Gow;Marcos García;Carolina Scarton;M. Idiart;Aline Villavicencio

文献摘要

参考文献

被引文献

相似文献

本文提出了多语言惯用性检测和句子嵌入的共享任务,该任务由两个子任务组成:(a)旨在识别句子是否包含惯用表达的二元分类任务,以及(b)基于语义文本相似性的任务,该任务要求模型在上下文中充分表示潜在的惯用表达。每个子任务都包含有关训练数据量的不同设置。除了任务描述之外,本文还介绍了英语、葡萄牙语和加利西亚语的数据集及其注释过程、评估指标以及参与者系统及其结果的摘要。该任务有近 100 名注册参与者被组织成 25 个团队,在实践和评估阶段分别提交了 650 多份和 150 份意见书。
This paper presents the shared task on Multilingual Idiomaticity Detection and Sentence Embedding, which consists of two subtasks: (a) a binary classification task aimed at identifying whether a sentence contains an idiomatic expression, and (b) a task based on semantic text similarity which requires the model to adequately represent potentially idiomatic expressions in context. Each subtask includes different settings regarding the amount of training data. Besides the task description, this paper introduces the datasets in English, Portuguese, and Galician and their annotation procedure, the evaluation metrics, and a summary of the participant systems and their results. The task had close to 100 registered participants organised into twenty five teams making over 650 and 150 submissions in the practice and evaluation phases respectively.
探索向量空间模型中的惯用法
DOI: --
发表时间: 2021
期刊: --
影响因子: --
作者:
Garcia, M
通讯作者: Garcia, M