Deep Learning Algorithms based Voiceprint Recognition System in Noisy Environment

Deep Learning Algorithms based Voiceprint Recognition System in Noisy Environment
复制标题

基于深度学习算法的噪声环境声纹识别系统

DOI:
--
复制
发表时间:
2021
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
S. Aliesawi
S. Aliesawi
中科院分区:
--
文献类型:
--
作者:
Hajer Y. Khdier;Wesam M. Jasim;S. Aliesawi

文献摘要

被引文献

相似文献

声纹识别(VPR)是一种利用用户声音特征来确定用户所谓身份的机制,该技术是世界上最有用和最常见的生物识别技术之一,特别是在与安全相关的领域。这些可用于身份验证、监控、说话人的法医识别以及各种相关活动。在这项工作中,尝试创建一个使用卷积神经网络(CNN)识别人类说话者身份的系统。这项工作使用了两种方法:MFCC-CNN 和 RW-CNN。第一种方法是使用MFCC的标准方法,使用音频中的特征,这些特征将被输入到CNN中执行处理。训练 CNN 将以图片形式输入,然后开始通过所提出的 CNN 进行训练。第二种方法,RW-CNN,与第一种方法步骤相同,但不经过直接进入CNN的MFCC阶段。其中,两种方法都使用相同的CNN结构。在这项工作中,RW-CNN 和 MFCC-CNN 的准确率都达到了 96%。两种方法的结果相似,无论有噪声还是没有噪声,但性能参差不齐。该系统可以深度学习大量的人声,具有高精度和最少的流程要求。
Voiceprint Recognition (VPR) is the mechanism by which a user’s so-called identity is determined using characteristics taken from their voice, where this-technique is one of the world’s most useful and common biometric recognition techniques particularly the fields-relevant to security. These can be used for authentication, monitoring, forensic identification of speakers, and a variety of related activities. In this work, an attempt is applied to create a system that recognizes human speaker identity using Convolutional Neural Network (CNN). Two methods are used in this work which are MFCC-CNN and RW-CNN. The first method is standard method using MFCC, to use the features in the audio, where these features are will be entered into CNN to perform a process. The training CNN will take input as a picture and then the process of training via the proposed CNN is beginning. The second method, RW-CNN, the same steps as the first method, but without going through the MFCC phases where direct entry to CNN. In which, the same CNN structure was used in both methods. In this work, a 96% accuracy gained for both RW-CNN and MFCC-CNN. Both methods are similar in their results, either with or without noise, but the performance is mixed. This system can deep learn a large amount of human voices with high accuracy and minimum processes requirement.