Word Discovery in Visually Grounded, Self-Supervised Speech Models

Word Discovery in Visually Grounded, Self-Supervised Speech Models
复制标题

基于视觉的自我监督语音模型中的单词发现

DOI:
10.48550/arxiv.2203.15081
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
David F. Harwath
David F. Harwath
中科院分区:
--
文献类型:
--
作者:
Puyuan Peng;David F. Harwath

文献摘要

参考文献

被引文献

相似文献

我们提出了一种视觉接地口语术语发现的方法。在训练HuBERT或wav2vec2.0模型将语音字幕与自然图像相关联后,我们发现,在模型的自我注意力头部中出现了强大的分词和聚类能力。我们的实验表明,这种能力是不存在的基础HuBERT和wav2vec2.0模型在几乎相同的程度上,这表明视觉接地任务是一个关键组成部分的单词发现能力,我们观察。我们还评估了我们的方法对七叶树词分割和ZeroSpeech口语术语发现任务,我们在几个指标上的表现与目前公布的方法相当或更好。代码和型号重量可在https://github.com/jasonppy/word-discovery上获得。
We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that powerful word segmentation and clustering capability emerges within the model's self-attention heads. Our experiments reveal that this ability is not present to nearly the same extent in the base HuBERT and wav2vec2.0 models, suggesting that the visual grounding task is a crucial component of the word discovery capability we observe. We also evaluate our method on the Buckeye word segmentation and ZeroSpeech spoken term discovery tasks, where we perform on par with or better than currently published methods on several metrics. Code and model weights are available at https://github.com/jasonppy/word-discovery.
DOI: 10.1109/asru51503.2021.9688093
发表时间: 2021-07
期刊: 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子: --
作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
通讯作者: Ankita Pasad;Ju-Chieh Chou;Karen Livescu