Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information

Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information
复制标题

沉默比言语更甜蜜:使用沉默存储说话者信息的自监督模型

DOI:
--
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
Hung
Hung
中科院分区:
--
文献类型:
--
作者:
Chiyu Feng;Po;Hung

文献摘要

参考文献

被引文献

相似文献

自我监督学习(SSL)最近取得了很大的进步。SSL语音模型在广泛的下游任务上实现了不错的性能,这表明它们从语音中提取了不同方面的信息。然而,SSL模型如何在不干扰的情况下以隐藏表示存储各种信息仍然知之甚少。以最近成功的SSL模型HuBERT为例,我们探讨了SSL模型如何在表示中处理和存储说话人信息。我们发现,HuBERT存储扬声器信息的表示,其位置对应于波形中的沉默。有几个证据。(1)我们发现,波形中含有较多无声部分的话语具有较好的说话人识别(SID)准确性。(2)如果我们使用整个话语进行SID,沉默部分总是对SID任务贡献更大。(3)如果我们只使用话语的一部分来表示SID,沉默部分比其他部分具有更高的准确性。我们的研究结果不仅有助于更好地理解SSL模型,而且还提高了性能。通过简单地在原始波形中添加静音,HuBERT将其SID的准确性提高了近2%。
Self-Supervised Learning (SSL) has made great strides recently. SSL speech models achieve decent performance on a wide range of downstream tasks, suggesting that they extract different aspects of information from speech. However, how SSL models store various information in hidden representations without interfering is still poorly understood. Taking the recently successful SSL model, HuBERT, as an example, we explore how the SSL model processes and stores speaker information in the representation. We found that HuBERT stores speaker information in representations whose positions correspond to silences in a waveform. There are several pieces of evidence. (1) We find that the utterances with more silent parts in the waveforms have better Speaker Identification (SID) accuracy. (2) If we use the whole utterances for SID, the silence part always contributes more to the SID task. (3) If we only use the representation of a part of the utterance for SID, the silenced part has higher accuracy than the other parts. Our findings not only contribute to a better understanding of SSL models but also improve performance. By simply adding silence to the original waveform, HuBERT improved its accuracy on SID by nearly 2%.
DOI: 10.1109/asru51503.2021.9688093
发表时间: 2021-07
期刊: 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子: --
作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
通讯作者: Ankita Pasad;Ju-Chieh Chou;Karen Livescu
DOI: 10.18653/v1/d19-1445
发表时间: 2019-08
期刊: ArXiv
影响因子: --
作者:
Olga Kovaleva;Alexey Romanov;Anna Rogers;Anna Rumshisky
通讯作者: Olga Kovaleva;Alexey Romanov;Anna Rogers;Anna Rumshisky