Augmenting Data for Sarcasm Detection with Unlabeled Conversation Context

Augmenting Data for Sarcasm Detection with Unlabeled Conversation Context
复制标题

使用未标记的对话上下文增强讽刺检测数据

DOI:
10.18653/v1/2020.figlang-1.2
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
Gunhee Kim
Gunhee Kim
中科院分区:
--
文献类型:
--
作者:
Hankyol Lee;Youngjae Yu;Gunhee Kim

文献摘要

被引文献

相似文献

我们提出了一种新的数据增强技术,CRA(上下文响应增强),它利用会话上下文来生成有意义的样本进行训练。我们还通过改变模型的输入输出格式,使其能够有效地处理不同的上下文长度,从而缓解了上下文长度不平衡的问题。具体来说,我们提出的模型,使用所提出的数据增强技术进行训练,参与了FigLang2020的讽刺检测任务,在Reddit和Twitter数据集中都获得了最佳性能。
We present a novel data augmentation technique, CRA (Contextual Response Augmentation), which utilizes conversational context to generate meaningful samples for training. We also mitigate the issues regarding unbalanced context lengths by changing the input output format of the model such that it can deal with varying context lengths effectively. Specifically, our proposed model, trained with the proposed data augmentation technique, participated in the sarcasm detection task of FigLang2020, have won and achieves the best performance in both Reddit and Twitter datasets.