FLEURS: FEW-Shot Learning Evaluation of Universal Representations of Speech

FLEURS: FEW-Shot Learning Evaluation of Universal Representations of Speech
复制标题

FLEURS:语音通用表示的 FEW-Shot 学习评估

DOI:
10.1109/slt54892.2023.10023141
复制
发表时间:
2022
期刊:
2022 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Ankur Bapna
Ankur Bapna
中科院分区:
--
文献类型:
--
作者:
Alexis Conneau;Min Ma;Simran Khanuja;Yu Zhang;Vera Axelrod;Siddharth Dalmia;Jason Riesa;Clara Rivera;Ankur Bapna

文献摘要

参考文献

被引文献

相似文献

我们介绍 FLEURS,即通用语音表示的少样本学习评估基准。 FLEURS 是基于机器翻译 FLoRes-101 基准构建的 102 种语言的 n 路并行语音数据集,每种语言大约有 12 小时的语音监督。 FLEURS 可用于各种语音任务,包括自动语音识别 (ASR)、语音语言识别 (Speech LangID)、语音文本检索。在本文中,我们为基于多语言预训练模型(例如纯语音 w2v-BERT [1] 和语音文本多模态 mSLAM [2])的任务提供了基线。 FLEURS 的目标是使语音技术能够支持更多语言,并促进低资源语音理解方面的研究。1.
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Speech-Text Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like speech-only w2v-BERT [1] and speech-text multimodal mSLAM [2]. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.1.
DOI: 10.1109/icassp40776.2020.9054362
发表时间: 2020-02
期刊: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze
通讯作者: Xinjian Li;Siddharth Dalmia;Juncheng Billy Li;Matthew Russell Lee;Patrick Littell;Jiali Yao;Antonios Anastasopoulos;David R. Mortensen;Graham Neubig;A. Black;Florian Metze