Self-organizing network for variable clustering

Self-organizing network for variable clustering
复制标题

DOI:
10.1007/s10479-017-2442-2
复制
发表时间:
2018-04-01
影响因子:
4.8
通讯作者:
Yang, Hui
Yang, Hui
中科院分区:
管理学3区
文献类型:
--
作者:
Liu, Gang;Yang, Hui

文献摘要

被引文献

相似文献

先进的传感技术和物联网带来了大数据,这为数据驱动的知识发现提供了前所未有的机会。然而,在大数据中涉及大量变量(或预测因子、特征)是常见的。变量之间复杂的相互依赖结构对传统的预测建模框架提出了重大挑战。本文提出了一种新的自组织网络的方法来描述变量之间的相互关系,并将它们聚类到同质子组的预测建模。具体来说,我们开发了一种新的方法,即非线性耦合分析来衡量变量对变量的相互依赖结构。此外,每个变量被表示为复杂网络中的节点。非线性耦合力移动这些节点,以获得网络的自组织拓扑结构。因此,变量被聚集到子网络社区中。仿真实验结果表明,该方法不仅优于传统的变量聚类算法,如层次聚类和斜主成分分析,而且能有效地识别变量之间的相互依赖结构,进一步提高预测建模的性能.此外,真实世界的案例研究表明,所提出的方法产生的平均灵敏度为96.80%,平均特异性为92.62%,在识别心肌梗死使用稀疏参数的心电向量表示模型。提出的自组织网络的新思想普遍适用于许多学科的预测建模,涉及大量的高度冗余的变量。
Advanced sensing and internet of things bring the big data, which provides an unprecedented opportunity for data-driven knowledge discovery. However, it is common that a large number of variables (or predictors, features) are involved in the big data. Complex interdependence structures among variables pose significant challenges on the traditional framework of predictive modeling. This paper presents a new methodology of self-organizing network to characterize the interrelationships among variables and cluster them into homogeneous subgroups for predictive modeling. Specifically, we develop a new approach, namely nonlinear coupling analysis to measure variable-to-variable interdependence structures. Further, each variable is represented as a node in the complex network. Nonlinear-coupling forces move these nodes to derive a self-organizing topology of the network. As such, variables are clustered into sub-network communities. Results of simulation experiments demonstrate that the proposed method not only outperforms traditional variable clustering algorithms such as hierarchical clustering and oblique principal component analysis, but also effectively identifies interdependent structures among variables and further improves the performance of predictive modeling. Additionally, real-world case study shows that the proposed method yields an average sensitivity of 96.80% and an average specificity of 92.62% in the identification of myocardial infarctions using sparse parameters of vectorcardiogram representation models. The proposed new idea of self-organizing network is generally applicable for predictive modeling in many disciplines that involve a large number of highly-redundant variables.