Personality trait estimation in group discussions using multimodal analysis and speaker embedding

Personality trait estimation in group discussions using multimodal analysis and speaker embedding
复制标题

DOI:
10.1007/s12193-023-00401-0
复制
发表时间:
2023-02
影响因子:
2.9
通讯作者:
Candy Olivia Mawalim;S. Okada;Y. Nakano;M. Unoki
Candy Olivia Mawalim;S. Okada;Y. Nakano;M. Unoki
中科院分区:
计算机科学3区
文献类型:
--
作者:
Candy Olivia Mawalim;S. Okada;Y. Nakano;M. Unoki

文献摘要

相似文献

人格特质的自动估计是许多人机接口(HCI)应用的基础。本文利用最新的说话人个性特征,即身份向量(i-vector)说话人嵌入,通过多模态分析和迁移学习,改进了小组讨论中的大五人格特质估计。通过研究用于估计的有效和鲁棒的多模态特征来进行实验,两个组讨论数据集,即,多模式任务导向小组讨论(MATRICS)(日语)和紧急领导力(ELEA)(欧洲语言)语料库。随后,通过使用留一人交叉验证(LOPCV)和消融测试进行评估,以比较每种模式的有效性。总体结果表明,说话人相关特征,例如,i向量的引入,有效地提高了大五人格特质估计的预测精度。此外,实验结果表明,在两个语料库中,与音频相关的特征都是最突出的特征。
The automatic estimation of personality traits is essential for many human–computer interface (HCI) applications. This paper focused on improving Big Five personality trait estimation in group discussions via multimodal analysis and transfer learning with the state-of-the-art speaker individuality feature, namely, the identity vector (i-vector) speaker embedding. The experiments were carried out by investigating the effective and robust multimodal features for estimation with two group discussion datasets, i.e., the Multimodal Task-Oriented Group Discussion (MATRICS) (in Japanese) and Emergent Leadership (ELEA) (in European languages) corpora. Subsequently, the evaluation was conducted by using leave-one-person-out cross-validation (LOPCV) and ablation tests to compare the effectiveness of each modality. The overall results showed that the speaker-dependent features, e.g., the i-vector, effectively improved the prediction accuracy of Big Five personality trait estimation. In addition, the experimental results showed that audio-related features were the most prominent features in both corpora.