Spoken Language Identification using ConvNets

Spoken Language Identification using ConvNets
复制标题

DOI:
10.1007/978-3-030-34255-5_17
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Sarthak;Shikhar Shukla;Govind Mittal
Sarthak;Shikhar Shukla;Govind Mittal
中科院分区:
其他
文献类型:
--
作者:
Sarthak;Shikhar Shukla;Govind Mittal

文献摘要

被引文献

相似文献

语言识别 (LI) 是多种语音处理系统中重要的第一步。随着基于语音的助手数量的不断增加,语音智能已经成为一个广泛研究的领域。为了解决识别语言的问题,我们可以采用仅存在某种语言的语音的隐式方法,或者采用文本及其相应转录本的显式方法。由于缺乏转录数据,本文重点关注隐式方法。本文对现有模型进行了基准测试,并提出了一种新的基于注意力的语言识别模型,该模型使用 log-Mel 频谱图图像作为输入。我们还展示了原始波形作为 LI 任务神经网络模型特征的有效性。为了模型的训练和评估,我们从 VoxForge 数据集中获得了 6 种语言(英语、法语、德语、西班牙语、俄语和意大利语),准确率达到 95.4%;分类为 4 种语言(英语、法语、德语、西班牙语),准确率达到 96.3%。这种方法可以进一步扩展以包含更多语言。
Language Identification (LI) is an important first step in several speech processing systems. With a growing number of voice-based assistants, speech LI has emerged as a widely researched field. To approach the problem of identifying languages, we can either adopt an implicit approach where only the speech for a language is present or an explicit one where text is available with its corresponding transcript. This paper focuses on an implicit approach due to the absence of transcriptive data. This paper benchmarks existing models and proposes a new attention based model for language identification which uses log-Mel spectrogram images as input. We also present the effectiveness of raw waveforms as features to neural network models for LI tasks. For training and evaluation of models, we classified six languages (English, French, German, Spanish, Russian and Italian) with an accuracy of 95.4% and four languages (English, French, German, Spanish) with an accuracy of 96.3% obtained from the VoxForge dataset. This approach can further be scaled to incorporate more languages.