TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus

TEDxSK and JumpSK: A New Slovak Speech Recognition Dedicated Corpus
复制标题

TEDxSK 和 JumpSK:新的斯洛伐克语语音识别专用语料库

DOI:
10.1515/jazcas-2017-0044
复制
发表时间:
2017
期刊:
Journal of Linguistics/Jazykovedný casopis
影响因子:
--
通讯作者:
Tomás Koctúr
Tomás Koctúr
中科院分区:
--
文献类型:
--
作者:
Ján Staš;D. Hládek;P. Viszlay;Tomás Koctúr

文献摘要

参考文献

被引文献

相似文献

本文描述了一个新的斯洛伐克语音识别专用语料库,该语料库是从TEDx Talks和Jump斯洛伐克Lessons中建立的。拟议的语音库包括220个讲座和讲座,总时长约为58小时。通过基于主成分分析的声学语音分割和两个互补的语音识别系统的自动语音转录,以无监督的方式自动生成标注语音库。评价数据包括50个人工标注的讲座和讲座,总时长约为12个小时,用于评价斯洛伐克语语音识别的质量。通过对TEDx演讲和Jump斯洛伐克演讲的无监督自动标注,我们获得了21.26%的新语音段,错误率约为9.44%,适合于对预先训练的声学模型进行再训练或改编。
Abstract This paper describes a new Slovak speech recognition dedicated corpus built from TEDx talks and Jump Slovakia lectures. The proposed speech database consists of 220 talks and lectures in total duration of about 58 hours. Annotated speech database was generated automatically in an unsupervised manner by using acoustic speech segmentation based on principal component analysis and automatic speech transcription using two complementary speech recognition systems. The evaluation data consisting of 50 manually annotated talks and lectures in total duration of about 12 hours, has been created for evaluation of the quality of Slovak speech recognition. By unsupervised automatic annotation of TEDx talks and Jump Slovakia lectures we have obtained 21.26% of new speech segments with approximately 9.44% word error rate, suitable for retraining or adaptation of acoustic models trained beforehand.
DOI: --
发表时间: 2004
期刊: In Proc. ICSLP 1
影响因子: --
作者:
K.Shitaoka;H.Nanjo;T.Kawahara
通讯作者: T.Kawahara
DOI: 10.21437/eurospeech.2001-396
发表时间: 2001-09
期刊: --
影响因子: --
作者:
Akinobu Lee;Tatsuya Kawahara;K. Shikano
通讯作者: Akinobu Lee;Tatsuya Kawahara;K. Shikano
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
Y. Akita;M. Watanabe;and T. Kawahara
通讯作者: and T. Kawahara