Robust HMM-based speech/music segmentation

Robust HMM-based speech/music segmentation
复制标题

基于 HMM 的鲁棒语音/音乐分割

DOI:
10.1109/icassp.2002.5743713
复制
发表时间:
2002
期刊:
2002 IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
H. Bourlard
H. Bourlard
中科院分区:
--
文献类型:
--
作者:
J. Ajmera;I. McCowan;H. Bourlard

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的方法对高性能的语音/音乐分割有关的广播新闻的自动转录的现实任务。在这里提出的方法中,本地概率密度函数(PDF)的估计干净的麦克风语音训练被用作一个通道模型的输出的熵和“动态”将被测量和集成随着时间的推移,通过一个2状态(语音和非语音)隐马尔可夫模型(HMM)与最小持续时间的限制。HMM的参数是使用EM算法以完全无监督的方式训练的。不同的实验,包括各种语音和音乐风格,以及语音和音乐信号的不同段持续时间(真实的数据分布,主要是语音,或主要是音乐),将说明该方法的鲁棒性,在每种情况下,实现了大于94%的帧级精度。
In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the local probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.