Augmenting pre-trained language models with audio feature embedding for argumentation mining in political debates

Augmenting pre-trained language models with audio feature embedding for argumentation mining in political debates
复制标题

DOI:
10.18653/v1/2023.findings-eacl.21
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Rafael Mestre;Stuart Middleton;Matt Ryan;Masood Gheasi;T. Norman;Jiatong Zhu
Rafael Mestre;Stuart Middleton;Matt Ryan;Masood Gheasi;T. Norman;Jiatong Zhu
中科院分区:
其他
文献类型:
--
作者:
Rafael Mestre;Stuart Middleton;Matt Ryan;Masood Gheasi;T. Norman;Jiatong Zhu

文献摘要

被引文献

相似文献

自然语言处理(NLP)任务中的多模态集成寻求利用包含在两种或两种以上模态中的互补信息,如文本、音频和视频。本文以论证挖掘(AM)任务为例,研究了音频特征与文本的集成。我们采用先前报道的数据集并呈现音频增强版本(Multimodal USElecDeb60To16数据集)。我们报告了基于BERT和GloVe嵌入的两个文本模型,一个音频模型(基于CNN和Bi-LSTM)和多模态组合在28,850个话语数据集上的性能。结果表明,当使用完整数据集时,多模态模型并不优于基于文本的模型。然而,我们表明音频功能在数据有限的完全监督场景中增加了价值。我们发现,当数据稀缺时(例如原始数据集的10%),多模态模型可以提高性能,而基于BERT的文本模型则会大大降低性能。最后,我们进行了人工生成声音的研究和消融研究,以探讨不同音频特征在音频模型中的重要性。
The integration of multimodality in natural language processing (NLP) tasks seeks to exploit the complementary information contained in two or more modalities, such as text, audio and video. This paper investigates the integration of often under-researched audio features with text, using the task of argumentation mining (AM) as a case study. We take a previously reported dataset and present an audio-enhanced version (the Multimodal USElecDeb60To16 dataset). We report the performance of two text models based on BERT and GloVe embeddings, one audio model (based on CNN and Bi-LSTM) and multimodal combinations, on a dataset of 28,850 utterances. The results show that multimodal models do not outperform text-based models when using the full dataset. However, we show that audio features add value in fully supervised scenarios with limited data. We find that when data is scarce (e.g. with 10% of the original dataset) multimodal models yield improved performance, whereas text models based on BERT considerably decrease performance. Finally, we conduct a study with artificially generated voices and an ablation study to investigate the importance of different audio features in the audio models.