OPTIMIZE WAV2VEC2S ARCHITECTURE FOR SMALL TRAINING SET THROUGH ANALYZING ITS PRE-TRAINED MODELS ATTENTION PATTERN.

OPTIMIZE WAV2VEC2S ARCHITECTURE FOR SMALL TRAINING SET THROUGH ANALYZING ITS PRE-TRAINED MODELS ATTENTION PATTERN.
复制标题

通过分析其预训练模型的注意力模式,优化小型训练集的WAV2VEC2S架构。

DOI:
10.1109/icassp43922.2022.9747831
复制
发表时间:
2022
期刊:
Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing. ICASSP (Conference)
影响因子:
--
通讯作者:
Dodge,HirokoH
Dodge,HirokoH
中科院分区:
--
文献类型:
--
作者:
Chen,Liu;Asgari,Meysam;Dodge,HirokoH

文献摘要

相似文献

基于transformer的自动语音识别(ASR)系统已经在大数据集的情况下取得了成功。但是,在医学研究中,我们必须为非典型人群(即患有言语障碍的学龄前儿童)创建ASR,并且训练数据集较小。为了提高小数据集上的训练效率,我们通过分析其预训练模型的块级注意模式来优化Wav 2 Vec 2.0(Transformer的变体)的架构。我们表明,块级模式可以作为缩小优化方向的指标。为了确保实验的可重复性,我们利用Librispeech-100-clean作为训练数据来模拟有限的数据条件。我们利用两种技术,本地注意力机制和跨块参数共享,与反直观的配置。我们的优化架构优于香草架构约1.8%的绝对字错误率(WER)的dev-clean和1.4%的test-clean。
Transformer-based automatic speech recognition (ASR) systems have shown their success in the presence of large datasets. But, in medical research, we have to create ASR for the non-typical population, i.e. pre-school children with speech disorders, with small training dataset. To increase training efficiency on small datasets, we optimize the architecture of Wav2Vec 2.0, a variation of Transformer, through analyzing its pre-trained model’s block-level attention pattern. We show that block-level patterns can serve as an indicator for narrowing down the optimization direction. To ensure the reproducibility of our experiments, we leverage Librispeech-100-clean as training data to simulate the limited data condition. We leverage two techniques, local attention mechanism and cross-block parameter sharing, with counter-intuitive configurations. Our optimized architecture outperforms the vanilla architecture about 1.8% absolute word error rate (WER) on dev-clean and 1.4% on test-clean.