Speech recognition based on Itakura-Saito divergence and dynamics/sparseness constraints from mixed sound of speech and music by non-negative matrix factorization

Speech recognition based on Itakura-Saito divergence and dynamics/sparseness constraints from mixed sound of speech and music by non-negative matrix factorization
复制标题

DOI:
10.21437/interspeech.2014-160
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Naoaki Hashimoto;Shoichi Nakano;Kazumasa Yamamoto;S. Nakagawa
Naoaki Hashimoto;Shoichi Nakano;Kazumasa Yamamoto;S. Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Naoaki Hashimoto;Shoichi Nakano;Kazumasa Yamamoto;S. Nakagawa

文献摘要

相似文献

我们考虑了一种基于非负矩阵分解(NMF)的混合声音语音识别方法,该方法仅去除音乐,该混合声音由语音和音乐组成。我们使用Itakura-Saito散度代替Kullback-Leibler散度来比较代价函数,以及权重矩阵的动态和稀疏约束,以改善语音识别。对于使用匹配条件模型的孤立词识别,我们将单词错误率相对于没有删除音乐的情况降低了52:1%(平均从69.3%降低到85.3%)。索引术语:语音识别,混合声音,音乐去除,矢量量化,非负矩阵分解
We considered a speech recognition method for mixed sound, which is composed of both speech and music, that only removes music based on non-negative matrix factorization (NMF). We used Itakura-Saito divergence instead of Kullback-Leibler divergence to compare the cost function, and the dynamics and sparseness constraints of a weight matrix to improve speech recognition. For isolated word recognition using the matched condition model, we reduced the word error rate of 52:1% relative from the case that didn’t remove music (on average, from 69.3% to 85.3%). Index Terms: speech recognition, mixed sound, music removal, vector quantization, non-negative matrix factorization