Optimization of Voice/Music Detection in Sound Data

Optimization of Voice/Music Detection in Sound Data
复制标题

声音数据中语音/音乐检测的优化

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
M. Sugiyama
M. Sugiyama
中科院分区:
--
文献类型:
--
作者:
Shin'ichi Takeuchi;M. Yamashita;T. Uchida;M. Sugiyama

文献摘要

被引文献

相似文献

自动语音/音乐段检测被期望用于各种应用。对于语音识别和听写的一般应用,需要输入语音进行识别,以自动检测和删除音乐片段。为了检测语音和音乐片段,其中声音数据包含语音和音乐,本文提出了加权块倒谱通量(BCF),并使用区分训练技术优化的权重向量。本文还讨论了频率轴加权在计算倒谱通量和BCF中的有效性。这里,通过修改LPC倒谱距离计算来执行频率轴加权。实验结果表明,原始BCF的检测错误率为11.56%,低频加权的加权BCF对封闭数据的错误率为9.08%,对开放数据的错误率为10.48%。这一结果表明,在BCF计算的语音和音乐之间的检测的时间和频率轴加权的有效性。
Automatic voice/music segment detection is expected for various applications. For the general applications of voice recognition and dictation, input voice for the recognition is needed to detect and remove music section automatically. In order to detect voice and music segments, where sound data contains both voice and music, this paper proposes weighted Block Cepstrum Flux (BCF) and optimizes the weight vector using discriminative training technique. This paper also discusses the effectiveness of the frequency axis weighting in calculating Cepstrum Flux and BCF. Here, frequency axis weighting is carried out by the modification of LPC Cepstrum distance calculation. The experimental results shows the detection error rate of the original BCF is 11.56% and the error rate of the weighted BCF with the low-frequency weighting for closed data is 9.08 %, and 10.48 % for open data. This rensult shows the effectiveness of both time and frequency axis weighting in BCF calculation for detection between voice and music.