Optimization of Voice/Music Detection in Sound Data
Optimization of Voice/Music Detection in Sound Data
复制标题
声音数据中语音/音乐检测的优化
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
M. Sugiyama
中科院分区:
文献类型:
--
作者:
Shin'ichi Takeuchi;M. Yamashita;T. Uchida;M. Sugiyama
Automatic voice/music segment detection is expected for various applications. For the general applications of voice recognition and dictation, input voice for the recognition is needed to detect and remove music section automatically. In order to detect voice and music segments, where sound data contains both voice and music, this paper proposes weighted Block Cepstrum Flux (BCF) and optimizes the weight vector using discriminative training technique. This paper also discusses the effectiveness of the frequency axis weighting in calculating Cepstrum Flux and BCF. Here, frequency axis weighting is carried out by the modification of LPC Cepstrum distance calculation. The experimental results shows the detection error rate of the original BCF is 11.56% and the error rate of the weighted BCF with the low-frequency weighting for closed data is 9.08 %, and 10.48 % for open data. This rensult shows the effectiveness of both time and frequency axis weighting in BCF calculation for detection between voice and music.