Unsupervised Instance Discriminative Learning for Depression Detection from Speech Signals.

Unsupervised Instance Discriminative Learning for Depression Detection from Speech Signals.
复制标题

DOI:
10.21437/interspeech.2022-10814
复制
发表时间:
2022
期刊:
Interspeech
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

重度抑郁症(MDD)是一种影响数百万人的严重疾病,尽早诊断这种疾病至关重要。从语音信号中检测抑郁症对医生有很大的帮助,并且可以在没有任何侵入性程序的情况下完成。由于相关的标记数据是稀缺的,我们提出了一种改进的实例判别学习(IDL)方法,一种无监督的预训练技术,以提取增量不变和实例扩散嵌入。在学习增强不变嵌入方面,研究了语音的各种数据增强方法,时间掩蔽产生最好的性能。为了学习实例扩展嵌入,我们探索了对训练批次的实例进行采样的方法(基于不同说话者的随机采样)。结果发现,不同的扬声器为基础的采样提供了更好的性能比随机的,我们假设,这一结果是因为相关的扬声器信息被保存在嵌入。此外,我们提出了一种新的采样策略,伪实例为基础的采样(PIS),基于聚类算法,以提高扩展特性的嵌入。使用DepAudioNet在DAIC-WOZ(英语)和CONVERGE(普通话)数据集上进行实验,使用PIS在MDD检测中观察到统计学显著改善,p值分别为0.0015和0.05,相对于基线,没有预训练。
Major Depressive Disorder (MDD) is a severe illness that affects millions of people, and it is critical to diagnose this disorder as early as possible. Detecting depression from voice signals can be of great help to physicians and can be done without any invasive procedure. Since relevant labelled data are scarce, we propose a modified Instance Discriminative Learning (IDL) method, an unsupervised pre-training technique, to extract augment-invariant and instance-spread-out embeddings. In terms of learning augment-invariant embeddings, various data augmentation methods for speech are investigated, and time-masking yields the best performance. To learn instance-spreadout embeddings, we explore methods for sampling instances for a training batch (distinct speaker-based and random sampling). It is found that the distinct speaker-based sampling provides better performance than the random one, and we hypothesize that this result is because relevant speaker information is preserved in the embedding. Additionally, we propose a novel sampling strategy, Pseudo Instance-based Sampling (PIS), based on clustering algorithms, to enhance spread-out characteristics of the embeddings. Experiments are conducted with DepAudioNet on DAIC-WOZ (English) and CONVERGE (Mandarin) datasets, and statistically significant improvements, with p-value 0.0015 and 0.05, respectively, are observed using PIS in the detection of MDD relative to the baseline without pre-training.