Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts

Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts
复制标题

DOI:
10.1109/taffc.2018.2885304
复制
发表时间:
2018-12
影响因子:
11.2
通讯作者:
Duc Anh Phan;Yuji Matsumoto;Hiroyuki Shindo
Duc Anh Phan;Yuji Matsumoto;Hiroyuki Shindo
中科院分区:
计算机科学2区
文献类型:
--
作者:
Duc Anh Phan;Yuji Matsumoto;Hiroyuki Shindo

文献摘要

相似文献

文本情感检测是计算语言学和情感计算研究中的一个挑战,因为它涉及到发现给定文本中表达的所有相关情感。当应用于对话记录时,它变得更加困难,因为我们需要对说话者之间的口语进行建模,同时牢记整个对话的上下文。在本文中,我们提出了一个半监督多标签的方法来预测的情绪从会话成绩单。该语料库包含从电影中提取的会话报价。其中一小部分是注释的,其余的用于无监督训练。我们使用word2vec词嵌入方法从语料库中构建情感词典,并将话语嵌入到矢量表示中。然后使用深度学习自动编码器来发现无监督数据的底层结构。我们在标记的训练数据上微调学习模型,并在测试集上测量其性能。实验结果表明,该方法是有效的,只是稍微落后于人类注释者。
Textual emotion detection is a challenge in computational linguistics and affective computing study as it involves the discovery of all associated emotions expressed within a given piece of text. It becomes an even more difficult problem when applied to conversation transcripts, as we need to model the spoken utterances between speakers, keeping in mind the context of the entire conversation. In this paper, we propose a semisupervised multilabel method of predicting emotions from conversation transcripts. The corpus contains conversational quotes extracted from movies. A small number of them are annotated, while the rest are used for unsupervised training. We use the word2vec word-embedding method to build an emotion lexicon from the corpus and to embed the utterances into vector representations. A deep-learning autoencoder is then used to discover the underlying structure of the unsupervised data. We fine-tune the learned model on labeled training data, and measure its performance on a test set. The experiment result suggests that the method is effective and is only slightly behind human annotators.