The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition
复制标题

马凯雷雷无线电语音语料库:用于自动语音识别的卢干达无线电语料库

DOI:
10.48550/arxiv.2206.09790
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Josh Meyer
Josh Meyer
中科院分区:
--
文献类型:
--
作者:
Jonathan Mukiibi;A. Katumba;J. Nakatumba‐Nabende;Ali Hussein;Josh Meyer

文献摘要

被引文献

相似文献

对于资源不足的语言来说,建立可用的无线电监测自动语音识别(ASR)系统是一项具有挑战性的任务,但在无线电是公共交流和讨论的主要媒介的社会中,这一点至关重要。联合国在乌干达的初步努力证明,了解被排除在社交媒体之外的农村人口的看法对国家规划是多么重要。然而,这些努力正受到缺乏转录语音数据集的挑战。在本文中,Makerere人工智能研究实验室发布了一个155小时的卢甘达无线电语音语料库。据我们所知,这是撒哈拉以南非洲地区第一个公开可用的无线电数据集。本文描述了语音语料库的开发,并介绍了使用Coqui STT工具包(一个开源语音识别工具包)的基线Luganda ASR性能结果。
Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the United Nations in Uganda have proved how understanding the perceptions of rural people who are excluded from social media is important in national planning. However, these efforts are being challenged by the absence of transcribed speech datasets. In this paper, The Makerere Artificial Intelligence research lab releases a Luganda radio speech corpus of 155 hours. To our knowledge, this is the first publicly available radio dataset in sub-Saharan Africa. The paper describes the development of the voice corpus and presents baseline Luganda ASR performance results using Coqui STT toolkit, an open-source speech recognition toolkit.