Word Discovery in Visually Grounded, Self-Supervised Speech Models
Word Discovery in Visually Grounded, Self-Supervised Speech Models
复制标题
基于视觉的自我监督语音模型中的单词发现
DOI:
10.48550/arxiv.2203.15081
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
David F. Harwath
中科院分区:
文献类型:
--
作者:
Puyuan Peng;David F. Harwath
We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that powerful word segmentation and clustering capability emerges within the model's self-attention heads. Our experiments reveal that this ability is not present to nearly the same extent in the base HuBERT and wav2vec2.0 models, suggesting that the visual grounding task is a crucial component of the word discovery capability we observe. We also evaluate our method on the Buckeye word segmentation and ZeroSpeech spoken term discovery tasks, where we perform on par with or better than currently published methods on several metrics. Code and model weights are available at https://github.com/jasonppy/word-discovery.
DOI:
10.1109/asru51503.2021.9688093
发表时间:
2021-07
期刊:
2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子:
--
作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
通讯作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu