Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts
Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts
复制标题
DOI:
10.1109/taffc.2018.2885304
复制
发表时间:
2018-12
影响因子:
11.2
通讯作者:
Duc Anh Phan;Yuji Matsumoto;Hiroyuki Shindo
中科院分区:
文献类型:
--
作者:
Duc Anh Phan;Yuji Matsumoto;Hiroyuki Shindo
Textual emotion detection is a challenge in computational linguistics and affective computing study as it involves the discovery of all associated emotions expressed within a given piece of text. It becomes an even more difficult problem when applied to conversation transcripts, as we need to model the spoken utterances between speakers, keeping in mind the context of the entire conversation. In this paper, we propose a semisupervised multilabel method of predicting emotions from conversation transcripts. The corpus contains conversational quotes extracted from movies. A small number of them are annotated, while the rest are used for unsupervised training. We use the word2vec word-embedding method to build an emotion lexicon from the corpus and to embed the utterances into vector representations. A deep-learning autoencoder is then used to discover the underlying structure of the unsupervised data. We fine-tune the learned model on labeled training data, and measure its performance on a test set. The experiment result suggests that the method is effective and is only slightly behind human annotators.