General Tensor Spectral Co-clustering for Higher-Order Data

General Tensor Spectral Co-clustering for Higher-Order Data
复制标题

DOI:
--
复制
发表时间:
2016-03
期刊:
--
影响因子:
--
通讯作者:
Tao Wu;Austin R. Benson;D. Gleich
Tao Wu;Austin R. Benson;D. Gleich
中科院分区:
其他
文献类型:
--
作者:
Tao Wu;Austin R. Benson;D. Gleich

文献摘要

被引文献

相似文献

谱聚类和联合聚类是数据分析中的众所周知的技术,最近的工作已经将谱聚类扩展到来自网络的平方、对称张量和超矩阵。我们开发了一种新的张量谱联合聚类方法,适用于任何非负张量的数据。应用我们的方法的结果是三模式张量的行,列和切片的同时聚类,并且该想法推广到任何数量的模式。我们设计的算法通过递归地将张量分成两部分来工作。我们还设计了一个新的措施,以了解张量中的每个集群的作用。我们的新算法和流水线证明在合成和现实世界的问题。在种植高阶集群结构的合成问题上,我们的方法是唯一一种在所有情况下都能可靠地识别种植结构的方法。在基于n元文本数据的张量上,我们识别停用词和语义独立集;在来自航空公司-机场多式联运网络的张量上,我们发现航空公司和机场的全球和区域合作集群;在来自电子邮件网络的张量上,我们识别日常垃圾邮件和集中主题集。
Spectral clustering and co-clustering are well-known techniques in data analysis, and recent work has extended spectral clustering to square, symmetric tensors and hypermatrices derived from a network. We develop a new tensor spectral co-clustering method that applies to any non-negative tensor of data. The result of applying our method is a simultaneous clustering of the rows, columns, and slices of a three-mode tensor, and the idea generalizes to any number of modes. The algorithm we design works by recursively bisecting the tensor into two pieces. We also design a new measure to understand the role of each cluster in the tensor. Our new algorithm and pipeline are demonstrated in both synthetic and real-world problems. On synthetic problems with a planted higher-order cluster structure, our method is the only one that can reliably identify the planted structure in all cases. On tensors based on n-gram text data, we identify stop-words and semantically independent sets; on tensors from an airline-airport multimodal network, we find worldwide and regional co-clusters of airlines and airports; and on tensors from an email network, we identify daily-spam and focused-topic sets.