What You Say or How You Say It? Depression Detection Through Joint Modeling of Linguistic and Acoustic Aspects of Speech

What You Say or How You Say It? Depression Detection Through Joint Modeling of Linguistic and Acoustic Aspects of Speech
复制标题

DOI:
10.1007/s12559-020-09808-3
复制
发表时间:
2021-02-24
影响因子:
5.4
通讯作者:
Vinciarelli, Alessandro
Vinciarelli, Alessandro
中科院分区:
计算机科学2区
文献类型:
--
作者:
Aloshban, Nujud;Esposito, Anna;Vinciarelli, Alessandro

文献摘要

被引文献

相似文献

抑郁症是最常见的心理健康问题之一。(根据最近的估计,它影响了世界上超过4%的人口。)本文表明,结合语音的语言和声学方面的分析,可以区分抑郁和非抑郁的说话人,准确率在80%以上。工作中使用的方法是基于为序列建模(双向长期-短期记忆网络)和多模式分析方法(后期融合、联合表示和门控多模式单元)设计的网络。这些实验是在59次访谈(大约4个小时的材料)的语料库上进行的,涉及29名被诊断为抑郁症的人和30名对照参与者。除了80%的准确率外,结果还表明,由于人们倾向于只通过一种模式来表现他们的状况,因此多模式方法比单模式方法表现得更好,这是单模式方法的多样性的来源。此外,实验表明,可以测量该方法的置信度,并自动识别测试数据中性能高于预定义阈值的子集。通过使用基于语音和语言的自动分析的不显眼和廉价的技术,可以有效地检测抑郁。
Depression is one of the most common mental health issues. (It affects more than 4% of the world's population, according to recent estimates.) This article shows that the joint analysis of linguistic and acoustic aspects of speech allows one to discriminate between depressed and nondepressed speakers with an accuracy above 80%. The approach used in the work is based on networks designed for sequence modeling (bidirectional Long-Short Term Memory networks) and multimodal analysis methodologies (late fusion, joint representation and gated multimodal units). The experiments were performed over a corpus of 59 interviews (roughly 4 hours of material) involving 29 individuals diagnosed with depression and 30 control participants. In addition to an accuracy of 80%, the results show that multimodal approaches perform better than unimodal ones owing to people's tendency to manifest their condition through one modality only, a source of diversity across unimodal approaches. In addition, the experiments show that it is possible to measure the "confidence" of the approach and automatically identify a subset of the test data in which the performance is above a predefined threshold. It is possible to effectively detect depression by using unobtrusive and inexpensive technologies based on the automatic analysis of speech and language.