Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification

Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification
复制标题

利用频谱增强进行语码转换口语识别

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Hemant Misra
Hemant Misra
中科院分区:
--
文献类型:
--
作者:
P. Rangan;Sundeep Teki;Hemant Misra

文献摘要

被引文献

相似文献

需要口语识别(LID)系统来识别给定音频样本中存在的语言,并且通常可以是许多语音处理相关任务(诸如自动语音识别(ASR))中的第一步。自动识别语音信号中的语言不仅在科学上很有趣,而且在印度这样的多语言国家也具有重要的实际意义。在许多印度城市,当人们相互交流时,可能会混合多达三种语言。这些可能包括该省的官方语言,印地语和英语(有时邻近省份的语言也可能在这些互动中混合)。这使得口语LID任务在印度背景下极具挑战性。虽然在印度语言的背景下已经实现了相当多的LID系统,但大多数这样的系统使用的是组织内部收集的小规模语音数据。在目前的工作中,我们对三种印度语言(古吉拉特语,泰卢固语和泰米尔语)与英语混合的代码进行口语LID。这项任务是由微软研究团队组织的,作为一个口头LID挑战。在我们的工作中,我们修改了通常的频谱增强方法,并提出了一个语言掩码,区分语言ID对,这导致了噪声鲁棒的口语LID系统。所提出的方法相对于微软提出的基线系统,在挑战中建议的两个共享任务的三种语言对上,LID准确性相对提高了约3-5%。
Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech recognition (ASR). Automatic identification of the languages present in a speech signal is not only scientifically interesting, but also of practical importance in a multilingual country such as India. In many of the Indian cities, when people interact with each other, as many as three languages may get mixed. These may include the official language of that province, Hindi and English (at times the languages of the neighboring provinces may also get mixed during these interactions). This makes the spoken LID task extremely challenging in Indian context. While quite a few LID systems in the context of Indian languages have been implemented, most such systems have used small scale speech data collected internally within an organization. In the current work, we perform spoken LID on three Indian languages (Gujarati, Telugu, and Tamil) code-mixed with English. This task was organized by the Microsoft research team as a spoken LID challenge. In our work, we modify the usual spectral augmentation approach and propose a language mask that discriminates the language ID pairs, which leads to a noise robust spoken LID system. The proposed method gives a relative improvement of approximately 3-5% in the LID accuracy over a baseline system proposed by Microsoft on the three language pairs for two shared tasks suggested in the challenge.