Singing Voice Detection Using Multi-Feature Deep Fusion with CNN

Singing Voice Detection Using Multi-Feature Deep Fusion with CNN
复制标题

DOI:
10.1007/978-981-15-2756-2_4
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Xulong Zhang;Shengchen Li;Zijin Li;Shizhe Chen;Yongwei Gao;Wei Li-
Xulong Zhang;Shengchen Li;Zijin Li;Shizhe Chen;Yongwei Gao;Wei Li-
中科院分区:
其他
文献类型:
--
作者:
Xulong Zhang;Shengchen Li;Zijin Li;Shizhe Chen;Yongwei Gao;Wei Li-

文献摘要

被引文献

相似文献

歌声检测的问题是将歌曲分割成有声部分和非有声部分。通常使用的方法通常在一组基于帧的特征上训练模型,然后通过模型预测未知帧。然而,多维特征通常对于每个帧连接在一起,很少考虑空间信息。为此,提出了一种基于卷积神经网络(CNN)的多特征维度深度融合方法。对每帧图像的特征维数进行一维卷积,得到的高层特征可用于直接的二值分类。该方法的性能与公共数据集上的最新方法相当。
The problem of singing voice detection is to segment a song into vocal and non-vocal parts. Commonly used methods usually train a model on a set of frame-based features and then predict the unknown frames by the model. However, the multi-dimensional features are usually concatenated together for each frame, with little consideration of spatial information. Hence, a deep fusion method of the Multi-feature dimensions with Convolution Neural Networks (CNN) is proposed. A one dimension convolution is made on feature dimensions for each frames, then the high-level features obtained can be used for a direct binary classification. The performance of the proposed method is on par with the state-of-art methods on public dataset.