Fine-tuning pre-trained models for Automatic Speech Recognition, experiments on a fieldwork corpus of Japhug (Trans-Himalayan family)
Fine-tuning pre-trained models for Automatic Speech Recognition, experiments on a fieldwork corpus of Japhug (Trans-Himalayan family)
复制标题
微调自动语音识别的预训练模型,在 Japhug(跨喜马拉雅家族)的实地工作语料库上进行实验
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Maxime Fily
中科院分区:
文献类型:
--
作者:
Severine Guillaume;Guillaume Wisniewski;Cécile Macaire;Guillaume Jacques;Alexis Michaud;Benjamin Galliot;Maximin Coavoux;Solange Rossato;Minh;Maxime Fily
This is a report on results obtained in the development of speech recognition tools intended to support linguistic documentation efforts. The test case is an extensive fieldwork corpus of Japhug, an endangered language of the Trans-Himalayan (Sino-Tibetan) family. The goal is to reduce the transcription workload of field linguists. The method used is a deep learning approach based on the language-specific tuning of a generic pre-trained representation model, XLS-R, using a Transformer architecture. We note difficulties in implementation, in terms of learning stability. But this approach brings significant improvements nonetheless. The quality of phonemic transcription is improved over earlier experiments; and most significantly, the new approach allows for reaching the stage of automatic word recognition. Subjective evaluation of the tool by the author of the training data confirms the usefulness of this approach.
DOI:
10.21437/interspeech.2021-1970
发表时间:
2021-08
期刊:
--
影响因子:
--
作者:
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux
通讯作者:
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux