Non-uniform Speaker Disentanglement For Depression Detection From Raw Speech Signals.

Non-uniform Speaker Disentanglement For Depression Detection From Raw Speech Signals.
复制标题

DOI:
10.21437/interspeech.2023-2101
复制
发表时间:
2023-08
期刊:
Interspeech
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

虽然使用说话者身份特征(如说话者嵌入)的基于语音的抑郁症检测方法很受欢迎,但它们往往会损害患者的隐私。为了解决这个问题,我们提出了一个扬声器解纠缠的方法,利用对抗SID损失最大化的非均匀机制。这是通过在训练过程中改变模型不同层之间的对抗权重来实现的。我们发现,初始层的对抗权重越大,性能越好。我们使用ECAPA-TDNN模型的方法在DAIC-WoZ数据集上实现了0.7349的F1分数(比仅音频SOTA提高了3.7%),同时将说话人识别准确率降低了50%。我们的研究结果表明,通过语音信号识别抑郁症可以在不过分依赖说话者身份的情况下完成,为隐私保护的抑郁症检测方法铺平了道路。
While speech-based depression detection methods that use speaker-identity features, such as speaker embeddings, are popular, they often compromise patient privacy. To address this issue, we propose a speaker disentanglement method that utilizes a non-uniform mechanism of adversarial SID loss maximization. This is achieved by varying the adversarial weight between different layers of a model during training. We find that a greater adversarial weight for the initial layers leads to performance improvement. Our approach using the ECAPA-TDNN model achieves an F1-score of 0.7349 (a 3.7% improvement over audio-only SOTA) on the DAIC-WoZ dataset, while simultaneously reducing the speaker-identification accuracy by 50%. Our findings suggest that identifying depression through speech signals can be accomplished without placing undue reliance on a speaker’s identity, paving the way for privacy-preserving approaches of depression detection.