Fine-tuning pre-trained models for Automatic Speech Recognition, experiments on a fieldwork corpus of Japhug (Trans-Himalayan family)

Fine-tuning pre-trained models for Automatic Speech Recognition, experiments on a fieldwork corpus of Japhug (Trans-Himalayan family)
复制标题

微调自动语音识别的预训练模型,在 Japhug(跨喜马拉雅家族)的实地工作语料库上进行实验

DOI:
--
复制
发表时间:
2022
期刊:
COMPUTEL
影响因子:
--
通讯作者:
Maxime Fily
Maxime Fily
中科院分区:
--
文献类型:
--
作者:
Severine Guillaume;Guillaume Wisniewski;Cécile Macaire;Guillaume Jacques;Alexis Michaud;Benjamin Galliot;Maximin Coavoux;Solange Rossato;Minh;Maxime Fily

文献摘要

参考文献

被引文献

相似文献

这是一份关于开发语音识别工具以支持语言文件工作所取得成果的报告。测试案例是一个广泛的田野调查语料库,日本语是跨喜马拉雅(汉藏)语系的一种濒危语言。这样做的目的是减少现场语言学家的转录工作量。使用的方法是一种深度学习方法,基于使用Transformer体系结构的通用预训练表示模型XLS-R的特定于语言的调优。我们注意到在学习稳定性方面的执行困难。但尽管如此,这种方法还是带来了显著的改善。音素转录的质量比以前的实验有所提高;最重要的是,新方法允许达到自动单词识别的阶段。培训数据的作者对该工具的主观评价证实了这一方法的有效性。
This is a report on results obtained in the development of speech recognition tools intended to support linguistic documentation efforts. The test case is an extensive fieldwork corpus of Japhug, an endangered language of the Trans-Himalayan (Sino-Tibetan) family. The goal is to reduce the transcription workload of field linguists. The method used is a deep learning approach based on the language-specific tuning of a generic pre-trained representation model, XLS-R, using a Transformer architecture. We note difficulties in implementation, in terms of learning stability. But this approach brings significant improvements nonetheless. The quality of phonemic transcription is improved over earlier experiments; and most significantly, the new approach allows for reaching the stage of automatic word recognition. Subjective evaluation of the tool by the author of the training data confirms the usefulness of this approach.
DOI: 10.21437/interspeech.2021-1970
发表时间: 2021-08
期刊: --
影响因子: --
作者:
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux
通讯作者: Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux