tmn at SemEval-2023 Task 9: Multilingual Tweet Intimacy Detection Using XLM-T, Google Translate, and Ensemble Learning

tmn at SemEval-2023 Task 9: Multilingual Tweet Intimacy Detection Using XLM-T, Google Translate, and Ensemble Learning
复制标题

tmn at SemEval-2023 任务 9:使用 XLM-T、谷歌翻译和集成学习进行多语言推文亲密度检测

DOI:
10.18653/v1/2023.semeval-1.183
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Anna Glazkova
Anna Glazkova
中科院分区:
--
文献类型:
--
作者:
Anna Glazkova

文献摘要

参考文献

被引文献

相似文献

本文描述了一个基于transformer的系统设计SemEval-2023任务9:多语言推文亲密度分析。这项任务的目的是预测推特的亲密度,范围从1(一点也不亲密)到5(非常亲密)。比赛的官方培训包括六种语言(英语、西班牙语、意大利语、葡萄牙语、法语和中文)的推文。测试集包括给定的六种语言以及外部数据,其中四种语言未出现在训练集中(印地语,阿拉伯语,荷兰语和韩语)。我们提出了一个解决方案的基础上的XLM-T,一个多语言的RoberTa模型适用于Twitter域的合奏。为了提高看不见的语言的性能,每条推文都补充了英文翻译。我们探索了在微调中看到的语言的翻译数据与看不见的语言相比的有效性,并估计了在基于transformer的模型中使用翻译数据的策略。我们的解决方案在排行榜上排名第四,同时在测试集上实现了0.5989的总体Pearson's r。建议的系统提高了0.088皮尔逊的r在所有45个提交的平均得分。
The paper describes a transformer-based system designed for SemEval-2023 Task 9: Multilingual Tweet Intimacy Analysis. The purpose of the task was to predict the intimacy of tweets in a range from 1 (not intimate at all) to 5 (very intimate). The official training set for the competition consisted of tweets in six languages (English, Spanish, Italian, Portuguese, French, and Chinese). The test set included the given six languages as well as external data with four languages not presented in the training set (Hindi, Arabic, Dutch, and Korean). We presented a solution based on an ensemble of XLM-T, a multilingual RoBERTa model adapted to the Twitter domain. To improve the performance on unseen languages, each tweet was supplemented by its English translation. We explored the effectiveness of translated data for the languages seen in fine-tuning compared to unseen languages and estimated strategies for using translated data in transformer-based models. Our solution ranked 4th on the leaderboard while achieving an overall Pearson’s r of 0.5989 over the test set. The proposed system improves up to 0.088 Pearson’s r over a score averaged across all 45 submissions.
SemEval-2023 任务 9:多语言推文亲密度分析
DOI: 10.18653/v1/2023.semeval-1.309
发表时间: 2023
期刊: Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023
影响因子: --
作者:
Pei, Jiaxin;Silva, Vítor;Bos, Maarten;Liu, Yozen;Neves, Leonardo;Jurgens, David;Barbieri, Francesco
通讯作者: Barbieri, Francesco
量化语言的亲密程度
DOI: 10.18653/v1/2020.emnlp-main.428
发表时间: 2020
期刊: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP
影响因子: --
作者:
Pei, Jiaxin;Jurgens, David
通讯作者: Jurgens, David