Acoustic Scene Clustering Using Joint Optimization of Deep Embedding Learning and Clustering Iteration

Acoustic Scene Clustering Using Joint Optimization of Deep Embedding Learning and Clustering Iteration
复制标题

使用深度嵌入学习和聚类迭代联合优化的声学场景聚类

DOI:
10.1109/tmm.2019.2947199
复制
发表时间:
2020-06-01
影响因子:
7.3
通讯作者:
He, Qianhua
He, Qianhua
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Yanxiong;Liu, Mingle;He, Qianhua

文献摘要

被引文献

相似文献

近年来,在音频信号处理领域对声场景分类进行了研究。相比之下,声场景聚类是一个新兴的问题,研究较少。声学场景聚类的目的是在不使用先验信息和训练分类器的情况下,将同一类声学场景的录音合并成一个单一的聚类。在本研究中,我们提出了一种结合特征学习和聚类迭代过程的声学场景聚类方法。在该方法中,学习到的特征是从深度卷积神经网络(CNN)中提取的深度嵌入,而聚类算法是聚类层次聚类(AHC)。我们建立了一个统一的损失函数来对这两个过程进行积分和优化。对各种特点和方法进行了比较。实验结果表明,该方法在归一化互信息和聚类精度方面优于其他无监督方法。此外,深度嵌入优于许多最先进的功能。
Recent efforts have been made on acoustic scene classification in the audio signal processing community. In contrast, few studies have been conducted on acoustic scene clustering, which is a newly emerging problem. Acoustic scene clustering aims at merging the audio recordings of the same class of acoustic scene into a single cluster without using prior information and training classifiers. In this study, we propose a method for acoustic scene clustering that jointly optimizes the procedures of feature learning and clustering iteration. In the proposed method, the learned feature is a deep embedding that is extracted from a deep convolutional neural network (CNN), while the clustering algorithm is the agglomerative hierarchical clustering (AHC). We formulate a unified loss function for integrating and optimizing these two procedures. Various features and methods are compared. The experimental results demonstrate that the proposed method outperforms other unsupervised methods in terms of the normalized mutual information and the clustering accuracy. In addition, the deep embedding outperforms many state-of-the-art features.