Compositional clustering: Applications to multi-label object recognition and speaker identification

Compositional clustering: Applications to multi-label object recognition and speaker identification
复制标题

DOI:
10.1016/j.patcog.2023.109829
复制
发表时间:
2021-09
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Zeqian Li;Xinlu He;J. Whitehill
Zeqian Li;Xinlu He;J. Whitehill
中科院分区:
其他
文献类型:
--
作者:
Zeqian Li;Xinlu He;J. Whitehill

文献摘要

相似文献

我们考虑一种新的聚类任务,其中集群可以有组成关系,例如,一个集群包含矩形的图像,一个包含圆形的图像,第三个(组成)集群包含两个对象的图像。与层次聚类中的父集群代表的子集群的属性的交集,我们的问题是寻找组成集群,代表的组成集群的属性的工会。该任务的动机是最近开发的少量学习和嵌入模型(Alfassy等人,2019年; Li等人,2021),可以区分标签集,而不仅仅是单个标签,分配给例子。我们提出了三个新的算法-组合亲和传播(CAP),组合k-均值(CKM)和贪婪的组合重新分配(GCR)-可以划分成连贯的组的例子,并推断它们之间的组合结构。我们展示了有前途的结果,相比流行的算法,如高斯混合,模糊c-均值,凝聚聚类,OmniGlot和LibriSpeech数据集。我们的工作已应用于开放世界的多标签对象识别和说话人识别与日记与同时从多个扬声器的语音。
We consider a novel clustering task in which clusters can have compositional relationships, eg, one cluster contains images of rectangles, one contains images of circles, and a third (compositional) cluster contains images with both objects. In contrast to hierarchical clustering in which a parent cluster represents the intersection of properties of the child clusters, our problem is about finding compositional clusters that represent the union of the properties of the constituent clusters. This task is motivated by recently developed few-shot learning and embedding models (Alfassy et al., 2019; Li et al., 2021) that can distinguish the label sets, not just the individual labels, assigned to the examples. We propose three new algorithms–Compositional Affinity Propagation (CAP), Compositional k-means (CKM), and Greedy Compositional Reassignment (GCR)–that can partition examples into coherent groups and infer the compositional structure among them. We show promising results, compared to popular algorithms such as Gaussian mixtures, Fuzzy c-means, and Agglomerative Clustering, on the OmniGlot and LibriSpeech datasets. Our work has applications to open-world multi-label object recognition and speaker identification & diarization with simultaneous speech from multiple speakers.