Importantaug: A Data Augmentation Agent for Speech

Importantaug: A Data Augmentation Agent for Speech
复制标题

DOI:
10.1109/icassp43922.2022.9747003
复制
发表时间:
2021-12
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
V. Trinh;Hassan Salami Kavaki;Michael Mandel
V. Trinh;Hassan Salami Kavaki;Michael Mandel
中科院分区:
其他
文献类型:
--
作者:
V. Trinh;Hassan Salami Kavaki;Michael Mandel

文献摘要

相似文献

我们介绍了ImportantAug,一种通过向语音的不重要区域而不是重要区域添加噪声来增强语音分类和识别模型的训练数据的技术。每个话语的重要性是由数据增强代理预测的,该数据增强代理经过训练以最大化它添加的噪声量,同时最小化它对识别性能的影响。我们的方法的有效性说明了版本2的谷歌语音命令(GSC)数据集。在标准GSC测试集上,与传统的噪声增强相比,它实现了23.3%的相对错误率降低,传统的噪声增强将噪声应用于语音,而不考虑它可能最有效的地方。与没有数据增强的基线相比,它还提供了25.4%的错误率降低。此外,建议的ImportantAug优于传统的噪声增强和基线上的两个测试集与额外的噪声添加。
We introduce ImportantAug, a technique to augment training data for speech classification and recognition models by adding noise to unimportant regions of the speech and not to important regions. Importance is predicted for each utterance by a data augmentation agent that is trained to maximize the amount of noise it adds while minimizing its impact on recognition performance. The effectiveness of our method is illustrated on version two of the Google Speech Commands (GSC) dataset. On the standard GSC test set, it achieves a 23.3% relative error rate reduction compared to conventional noise augmentation which applies noise to speech without regard to where it might be most effective. It also provides a 25.4% error rate reduction compared to a baseline without data augmentation. Additionally, the proposed ImportantAug outperforms the conventional noise augmentation and the baseline on two test sets with additional noise added.