Dealing with Distances and Transformations for Fuzzy C-Means Clustering of Compositional Data

Dealing with Distances and Transformations for Fuzzy C-Means Clustering of Compositional Data
复制标题

DOI:
10.1007/s00357-012-9105-4
复制
发表时间:
2012-07-01
影响因子:
2
通讯作者:
Soto, Jesus A.
Soto, Jesus A.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Palarea-Albaladejo, Javier;Antoni Martin-Fernandez, Josep;Soto, Jesus A.

文献摘要

被引文献

相似文献

聚类技术基于对象和聚类之间的差异或距离测量。本文重点研究单纯形空间,其元素(组合)受到非负性和常和约束。任何涉及成分的数据分析都应满足两个主要原则:尺度不变性和子成分一致性。在模糊聚类方法中,FCM算法广泛应用于各个领域,但在处理组合时表现不佳。在这里,在组合原理的基础上讨论了单纯形中不同相异性的充分性以及常见对数比变换的行为。因此,提出了一种有根据的 FCM 组合物聚类策略。理论研究结果附有数字证据,并提供了我们建议的详细说明。最后,使用聚类文献中已知的营养数据集来说明案例研究。
Clustering techniques are based upon a dissimilarity or distance measure between objects and clusters. This paper focuses on the simplex space, whose elements-compositions-are subject to non-negativity and constant-sum constraints. Any data analysis involving compositions should fulfill two main principles: scale invariance and subcompositional coherence. Among fuzzy clustering methods, the FCM algorithm is broadly applied in a variety of fields, but it is not well-behaved when dealing with compositions. Here, the adequacy of different dissimilarities in the simplex, together with the behavior of the common log-ratio transformations, is discussed in the basis of compositional principles. As a result, a well-founded strategy for FCM clustering of compositions is suggested. Theoretical findings are accompanied by numerical evidence, and a detailed account of our proposal is provided. Finally, a case study is illustrated using a nutritional data set known in the clustering literature.