HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
复制标题

DOI:
10.1109/taslp.2021.3122291
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Mohamed, Abdelrahman
Mohamed, Abdelrahman
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hsu, Wei-Ning;Bolte, Benjamin;Mohamed, Abdelrahman

文献摘要

被引文献

相似文献

用于语音表示学习的自监督方法受到三个独特问题的挑战:(1)每个输入话语中有多个声音单元,(2)在预训练阶段没有输入声音单元的词典,以及(3)声音单元具有可变长度,没有显式分割。为了解决这三个问题,我们提出了用于自监督语音表示学习的隐藏单元BERT(HuBERT)方法,该方法利用离线聚类步骤为BERT类预测损失提供对齐的目标标签。我们方法的一个关键因素是只在掩蔽区域上应用预测损失,这迫使模型在连续输入上学习组合的声学和语言模型。HuBERT主要依赖于无监督聚类步骤的一致性,而不是分配的聚类标签的内在质量。从一个简单的100个聚类的k-means老师开始,使用两次聚类迭代,HuBERT模型在Librispeech(960 h)和Libri-light(60,000 h)基准测试中匹配或改进了最先进的wav 2 vec 2.0性能,其中包括10 min,1 h,10 h,100 h和960 h微调子集。使用1B参数模型,HuBERT在更具挑战性的dev-other和test-other评估子集上显示出高达19%和13%的相对WER降低。(一)(二)
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation. To deal with these three problems, we propose the Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss. A key ingredient of our approach is applying the prediction loss over the masked regions only, which forces the model to learn a combined acoustic and language model over the continuous inputs. HuBERT relies primarily on the consistency of the unsupervised clustering step rather than the intrinsic quality of the assigned cluster labels. Starting with a simple k-means teacher of 100 clusters, and using two iterations of clustering, the HuBERT model either matches or improves upon the state-of-the-art wav2vec 2.0 performance on the Librispeech (960 h) and Libri-light (60,000 h) benchmarks with 10 min, 1 h, 10 h, 100 h, and 960 h fine-tuning subsets. Using a 1B parameter model, HuBERT shows up to 19% and 13% relative WER reduction on the more challenging dev-other and test-other evaluation subsets.(1)(2)