Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization Approach

Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization Approach
复制标题

DOI:
10.48550/arxiv.2305.19391
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Tri Nguyen;Shahana Ibrahim;Xiao Fu
Tri Nguyen;Shahana Ibrahim;Xiao Fu
中科院分区:
其他
文献类型:
--
作者:
Tri Nguyen;Shahana Ibrahim;Xiao Fu

文献摘要

相似文献

最新的深度学习和成对相似性的集成基于基于约束的聚类 - 即$ \ textit {深度约束聚类} $(DCC) - 已证明可以有效地将弱监督纳入大规模数据聚类:少于1%的对相似性注释通常可以显着提高聚类精度。但是,除了经验成功之外,对DCC缺乏了解。此外,许多DCC范式对注释噪声敏感,但是性能保证的嘈杂的DCC方法在很大程度上难以捉摸。这项工作首先深入研究了DCC最近出现的逻辑损失函数,并表征了其理论属性。我们的结果表明,逻辑DCC损失确保了在合理条件下数据成员资格的可识别性,这可能会揭示其在实践中的有效性。在这种理解的基础上,提出了基于几何因素分析的新损失函数,以抵抗嘈杂的注释。结果表明,即使在$ \ textit {norknown} $注释困惑下,数据成员资格仍然可以是$ \ textit {可证明} $在我们建议的学习标准下确定的。在多个数据集上测试了建议的方法,以验证我们的主张。
The recent integration of deep learning and pairwise similarity annotation-based constrained clustering -- i.e., $\textit{deep constrained clustering}$ (DCC) -- has proven effective for incorporating weak supervision into massive data clustering: Less than 1% of pair similarity annotations can often substantially enhance the clustering accuracy. However, beyond empirical successes, there is a lack of understanding of DCC. In addition, many DCC paradigms are sensitive to annotation noise, but performance-guaranteed noisy DCC methods have been largely elusive. This work first takes a deep look into a recently emerged logistic loss function of DCC, and characterizes its theoretical properties. Our result shows that the logistic DCC loss ensures the identifiability of data membership under reasonable conditions, which may shed light on its effectiveness in practice. Building upon this understanding, a new loss function based on geometric factor analysis is proposed to fend against noisy annotations. It is shown that even under $\textit{unknown}$ annotation confusions, the data membership can still be $\textit{provably}$ identified under our proposed learning criterion. The proposed approach is tested over multiple datasets to validate our claims.