Speaker recognition with hybrid features from a deep belief network

Speaker recognition with hybrid features from a deep belief network
复制标题

DOI:
10.1007/s00521-016-2501-7
复制
发表时间:
2016-08
影响因子:
6
通讯作者:
Hazrat Ali;S. Tran;Emmanouil Benetos;A. Garcez
Hazrat Ali;S. Tran;Emmanouil Benetos;A. Garcez
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hazrat Ali;S. Tran;Emmanouil Benetos;A. Garcez

文献摘要

被引文献

相似文献

在许多音频应用中,从音频数据中学习表示已经显示出优于手工特征(诸如梅尔频率倒谱系数(MFCC))的优势。在大多数的表示学习方法中,连接系统被用来从固定长度的数据中学习和提取潜在特征。在本文中,我们提出了一种联合收割机的学习功能和MFCC功能的说话人识别任务,它可以应用于不同长度的音频脚本。特别是,我们研究了使用来自不同层次的深度信念网络的特征将音频数据量化为音频单词计数的向量。这些向量代表不同长度的音频脚本,使它们更容易训练分类器。我们在实验中表明,从DBN功能在不同的层的混合产生的音频字计数向量给更好的性能比MFCC功能。我们还可以通过结合音频词计数向量和MFCC特征来实现进一步的改进。
Learning representation from audio data has shown advantages over the handcrafted features such as mel-frequency cepstral coefficients (MFCCs) in many audio applications. In most of the representation learning approaches, the connectionist systems have been used to learn and extract latent features from the fixed length data. In this paper, we propose an approach to combine the learned features and the MFCC features for speaker recognition task, which can be applied to audio scripts of different lengths. In particular, we study the use of features from different levels of deep belief network for quantizing the audio data into vectors of audio word counts. These vectors represent the audio scripts of different lengths that make them easier to train a classifier. We show in the experiment that the audio word count vectors generated from mixture of DBN features at different layers give better performance than the MFCC features. We also can achieve further improvement by combining the audio word count vector and the MFCC features.