Spoken language identification using the speechdat corpus
Spoken language identification using the speechdat corpus
复制标题
使用speechdat语料库进行口语识别
DOI:
--
复制
发表时间:
1998
期刊:
影响因子:
--
通讯作者:
I. Trancoso
中科院分区:
文献类型:
--
作者:
D. Caseiro;I. Trancoso
Current language identification systems vary significantly in their complexity. The systems that use higher level linguistic information have the best performance. Nevertheless, that information is hard to collect for each new language. The system presented in this paper is easily extendable to new languages because it uses very little linguistic information. In fact, the presented system needs only one language specific phone recogniser (in our case the Portuguese one), and is trained with speech from each of the other languages. With the SpeechDat-M corpus, with 6 European languages (English, French, German, Italian, Portuguese and Spanish) our system achieved an identification rate of 83.4% on 5-second utterances, this result shows an improvement of 5% over our previous version, mainly through the use of a neural network classifier. Both the baseline and the full system were implemented in realtime.