Improved Transcription and Speaker Identification System for Concurrent Speech in Bahasa Indonesia Using Recurrent Neural Network

Improved Transcription and Speaker Identification System for Concurrent Speech in Bahasa Indonesia Using Recurrent Neural Network
复制标题

DOI:
10.1109/access.2021.3077441
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Muhammad Bagus Andra;T. Usagawa
Muhammad Bagus Andra;T. Usagawa
中科院分区:
计算机科学3区
文献类型:
--
作者:
Muhammad Bagus Andra;T. Usagawa

文献摘要

相似文献

印度尼西亚语是最突出的低资源语言之一,在通信辅助技术方面仍然缺乏发展。本文提出了一种改进的系统,用于生成成绩单和识别扬声器从并发的语音在印度尼西亚的巴伊亚。所提出的方法适用于诸如在线会议和远程会议的情况。该系统结合强化学习(RL)模型和音高感知的语音分离来识别并发语音中的说话人。利用递归神经网络(RNN)生成文本转录本,然后通过外部语言模型和拼写校正模型对其进行改进。所提出的系统能够识别多达5个扬声器与不同程度的信心,并生成一个成绩单,为他们每个人更好的质量相比,其他方法时,用几个指标进行评估。实验结果表明,与基线方法相比,该方法在单说话人情况下的误识率更高,在语音不清晰情况下的误识率分别为16.59%、26.72%和31.50%。
Bahasa Indonesia is one of the most prominent low-resource Languages that still lack development in regards to communication-assisting technology. This paper proposes an improved system for generating transcript and identifying speakers from a concurrent speech in Bahasa Indonesia. The proposed method is applicable in a situation such as an online meeting and remote conference. The system combines Reinforced Learning (RL) Model with pitch-aware speech separation to identify the speakers in a concurrent speech. A Recurrent Neural Network (RNN) is utilized to generate the text transcript which is later improved by an external language model and spelling correction model. The proposed system was able to identify up to 5 speakers with a variable degree of confidence and generate a transcript for each of them with better quality compared to other methods when evaluated with several metrics. The result shows that the proposed method perform better compared to the baseline method, even in the single-speaker situation, and function in the simultaneous-speech situation, with an average Word Error Rate (WER) of 16.59% for two speakers, 26.72% for three speakers, and 31.50% for four speakers.