Multimodal Depression Severity Score Prediction Using Articulatory Coordination Features and Hierarchical Attention Based Text Embeddings

Multimodal Depression Severity Score Prediction Using Articulatory Coordination Features and Hierarchical Attention Based Text Embeddings
复制标题

DOI:
10.21437/interspeech.2022-11099
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Nadee Seneviratne;C. Espy-Wilson
Nadee Seneviratne;C. Espy-Wilson
中科院分区:
其他
文献类型:
--
作者:
Nadee Seneviratne;C. Espy-Wilson

文献摘要

相似文献

预测抑郁症严重程度的多模式方法是一个经过深入研究的问题。我们提出了一种多模态抑郁症严重程度评分预测系统,该系统使用从声道变量 (TV) 导出的发音协调特征 (ACF) 和从自动语音识别工具获得的文本转录,与单模态分类器相比,均方根误差得到改善(音频和文本分别为 14.8% 和 11%)。使用阶梯回归 (ST-R) 方法和基于 TV 的 ACF 来训练多级卷积循环神经网络。 ST-R 方法有助于更好地捕捉抑郁严重程度评分的准数字性质。使用分层注意力网络 (HAN) 架构训练文本模型。该多模态系统是通过将会话级音频模型和 HAN 文本模型的嵌入与包含语音信号定时测量的会话级辅助特征向量相结合而开发的。我们还表明,该模型相当好地跟踪了受试者抑郁的严重程度,并且我们分析了预测与真实分数存在显着偏差的案例的根本原因。
Multimodal approaches to predict depression severity is a highly researched problem. We present a multimodal depression severity score prediction system that uses articulatory co-ordination features (ACFs) derived from vocal tract variables (TVs) and text transcriptions obtained from an automatic speech recognition tool that yields improvements of the root mean squared errors compared to unimodal classifiers (14.8% and 11% for audio and text, respectively). A multi-stage convolutional recurrent neural network was trained using a staircase regression (ST-R) approach with the TV based ACFs. The ST-R approach helps to better capture the quasi-numerical nature of the depression severity scores. A text model is trained using the Hierarchical Attention Network (HAN) architecture. The multimodal system is developed by combining embeddings from the session-level audio model and the HAN text model with a session-level auxiliary feature vector containing timing measures of the speech signal. We also show that this model tracks the severity of depression for subjects reasonably well and we analyze the underlying reasons for the cases with significant deviations of the predictions from the ground-truth score.