Globalphone: a multilingual speech and text database developed at karlsruhe university

Globalphone: a multilingual speech and text database developed at karlsruhe university
复制标题

DOI:
10.21437/icslp.2002-151
复制
发表时间:
2002-09
期刊:
--
影响因子:
--
通讯作者:
Tanja Schultz
Tanja Schultz
中科院分区:
其他
文献类型:
--
作者:
Tanja Schultz

文献摘要

被引文献

相似文献

本文介绍了多语种数据库GlobalPhone的设计,收集和目前的状态,一个正在进行的项目,自1995年以来在卡尔斯鲁厄大学。GlobalPhone是一个高质量的语音和文本数据库,适用于多种语言的大词汇量语音识别系统的开发。它已经成功地应用于语言无关和语言自适应语音识别。GlobalPhone目前覆盖15种语言,包括阿拉伯语、中文(普通话和上海话)、克罗地亚语、捷克语、法语、德语、日语、韩语、葡萄牙语、俄语、西班牙语、瑞典语、泰米尔语和土耳其语。该语料库包含超过300小时的转录语音,由1500多名母语成年人使用,并将很快从ELRA提供。
This paper describes the design, collection, and current status of the multilingual database GlobalPhone, an ongoing project since 1995 at Karlsruhe University. GlobalPhone is a high-quality read speech and text database in a large variety of languages which is suitable for the development of large vocabulary speech recognition systems in many languages. It has already been successfully applied to language independent and language adaptive speech recognition. GlobalPhone currently covers 15 languages Arabic, Chinese (Mandarin and Shanghai), Croatian, Czech, French, German, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish. The corpus contains more than 300 hours of transcribed speech spoken by more than 1500 native, adult speakers and will soon be available from ELRA.