Spoken language identification using the speechdat corpus

Spoken language identification using the speechdat corpus
复制标题

使用speechdat语料库进行口语识别

DOI:
--
复制
发表时间:
1998
期刊:
ICSLP
影响因子:
--
通讯作者:
I. Trancoso
I. Trancoso
中科院分区:
--
文献类型:
--
作者:
D. Caseiro;I. Trancoso

文献摘要

被引文献

相似文献

当前的语言识别系统在其复杂性方面差异很大。使用更高级别语言信息的系统具有最佳性能。然而,这些信息很难为每一种新语言收集。由于该系统使用的语言信息很少,因此很容易扩展到新的语言。事实上,该系统只需要一个特定于语言的电话识别器(在我们的例子中是葡萄牙语的),并使用其他语言的语音进行训练。在SpeechDat-M语料库上,使用6种欧洲语言(英语、法语、德语、意大利语、葡萄牙语和西班牙语),我们的系统对5秒语音的识别率达到了83.4%,这一结果表明,主要是通过使用神经网络分类器,我们的系统比以前的版本提高了5%。基线和整个系统都是实时实施的。
Current language identification systems vary significantly in their complexity. The systems that use higher level linguistic information have the best performance. Nevertheless, that information is hard to collect for each new language. The system presented in this paper is easily extendable to new languages because it uses very little linguistic information. In fact, the presented system needs only one language specific phone recogniser (in our case the Portuguese one), and is trained with speech from each of the other languages. With the SpeechDat-M corpus, with 6 European languages (English, French, German, Italian, Portuguese and Spanish) our system achieved an identification rate of 83.4% on 5-second utterances, this result shows an improvement of 5% over our previous version, mainly through the use of a neural network classifier. Both the baseline and the full system were implemented in realtime.