Toward Knowledge-Driven Speech-Based Models of Depression: Leveraging Spectrotemporal Variations in Speech Vowels

Toward Knowledge-Driven Speech-Based Models of Depression: Leveraging Spectrotemporal Variations in Speech Vowels
复制标题

DOI:
10.1109/bhi56158.2022.9926939
复制
发表时间:
2022-09
期刊:
2022 IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI)
影响因子:
--
通讯作者:
Kexin Feng;Theodora Chaspari
Kexin Feng;Theodora Chaspari
中科院分区:
其他
文献类型:
--
作者:
Kexin Feng;Theodora Chaspari

文献摘要

相似文献

与抑郁症相关的精神发育迟滞与元音产生的有形差异有关。本文研究了一种知识驱动的机器学习(ML)方法,该方法在元音水平上整合语音的谱时信息来识别抑郁症。低级语音描述符由经过元音分类训练的卷积神经网络(CNN)学习。这些低级描述符的时间演变通过长短期记忆(LSTM)模型在话语内和话语间的高级建模,该模型做出最终的抑郁决策。修改后的版本的本地可解释的模型不可知的解释(LIME)进一步用于识别的影响,低级别的spectrotemporal元音变化的决定,并观察高级别的时间变化的抑郁症的可能性。所提出的方法优于在不整合基于元音的信息的情况下对语音中的频谱时间信息进行建模的基线,以及使用传统韵律和频谱时间特征训练的ML模型。进行的可解释性分析表明,对应于非元音段的频谱时间信息的重要性低于基于元音的信息。进一步检查有和没有抑郁症的参与者的高层次信息捕获的分段的决策的可解释性。这项工作的发现可以为知识驱动的可解释决策支持系统提供基础,这些系统可以帮助临床医生更好地了解语音数据中的细粒度时间变化,最终增强心理健康诊断和护理。
Psychomotor retardation associated with depression has been linked with tangible differences in vowel production. This paper investigates a knowledge-driven machine learning (ML) method that integrates spectrotemporal information of speech at the vowel-level to identify the depression. Low-level speech descriptors are learned by a convolutional neural network (CNN) that is trained for vowel classification. The temporal evolution of those low-level descriptors is modeled at the high-level within and across utterances via a long short-term memory (LSTM) model that takes the final depression decision. A modified version of the Local Interpretable Model-agnostic Explanations (LIME) is further used to identify the impact of the low-level spectrotemporal vowel variation on the decisions and observe the high-level temporal change of the depression likelihood. The proposed method outperforms baselines that model the spectrotemporal information in speech without integrating the vowel-based information, as well as ML models trained with conventional prosodic and spectrotemporal features. The conducted explainability analysis indicates that spectrotemporal information corresponding to non-vowel segments less important than the vowel-based information. Explainability of the high-level information capturing the segment-by-segment decisions is further inspected for participants with and without depression. The findings from this work can provide the foundation toward knowledge-driven interpretable decision-support systems that can assist clinicians to better understand fine-grain temporal changes in speech data, ultimately augmenting mental health diagnosis and care.