Automated depression analysis using convolutional neural networks from speech

Automated depression analysis using convolutional neural networks from speech
复制标题

DOI:
10.1016/j.jbi.2018.05.007
复制
发表时间:
2018-07-01
影响因子:
4.5
通讯作者:
Cao, Cui
Cao, Cui
中科院分区:
医学3区
文献类型:
--
作者:
He, Lang;Cao, Cui

文献摘要

被引文献

相似文献

为了帮助临床医生有效地诊断一个人的抑郁症的严重程度,情感计算界和人工智能领域对设计自动化系统表现出越来越大的兴趣。语音特征为抑郁症的诊断提供了有用的信息。然而,人工设计和领域知识对于特征的选择仍然很重要,这使得过程耗费人力和主观性。近年来,基于神经网络的深度学习特征在各个领域表现出了优于手工特征的性能。在本文中,为了克服上述困难,我们提出了一种手工制作和深度学习相结合的特征,可以有效地从语音中衡量抑郁症的严重程度。在该方法中,首先建立深度卷积神经网络(DCNN),从语谱图和原始语音波形中学习深度学习特征。然后,我们从光谱图中人工提取最先进的纹理描述符,称为中值稳健扩展局部二值模式(MRELBP)。为了捕捉人工特征和深度学习特征之间的互补信息,我们提出了联合微调层,将原始特征和谱图DCNN结合起来,以提高抑郁识别的性能。此外,为了解决小样本问题,提出了一种数据增强方法。在AVEC2013和AVEC2014抑郁症数据库上进行的实验表明,与最先进的基于音频的方法相比,我们的方法对于抑郁症的诊断是稳健和有效的。
To help clinicians to efficiently diagnose the severity of a person's depression, the affective computing community and the artificial intelligence field have shown a growing interest in designing automated systems. The speech features have useful information for the diagnosis of depression. However, manually designing and domain knowledge are still important for the selection of the feature, which makes the process labor consuming and subjective. In recent years, deep-learned features based on neural networks have shown superior performance to hand-crafted features in various areas. In this paper, to overcome the difficulties mentioned above, we propose a combination of hand-crafted and deep-learned features which can effectively measure the severity of depression from speech. In the proposed method, Deep Convolutional Neural Networks (DCNN) are firstly built to learn deep-learned features from spectrograms and raw speech waveforms. Then we manually extract the state-of-the-art texture descriptors named median robust extended local binary patterns (MRELBP) from spectrograms. To capture the complementary information within the hand-crafted features and deep-learned features, we propose joint fine-tuning layers to combine the raw and spectrogram DCNN to boost the depression recognition performance. Moreover, to address the problems with small samples, a data augmentation method was proposed. Experiments conducted on AVEC2013 and AVEC2014 depression databases show that our approach is robust and effective for the diagnosis of depression when compared to state-of-the-art audio-based methods.