TENALIGN: Joint Tensor Alignment and Coupled Factorization

TENALIGN: Joint Tensor Alignment and Coupled Factorization
复制标题

DOI:
10.1109/icdm54844.2022.00067
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Yu-Nuo Wu;Uday Singh Saini;Jia Chen;E. Papalexakis
Yu-Nuo Wu;Uday Singh Saini;Jia Chen;E. Papalexakis
中科院分区:
其他
文献类型:
--
作者:
Yu-Nuo Wu;Uday Singh Saini;Jia Chen;E. Papalexakis

文献摘要

被引文献

相似文献

表示为张量的多模态数据集通常共享它们的一些模式。然而,即使在耦合模式之间可能存在一对一(或可能部分)的对应关系,也可能不会给出这种对应关系/对准,特别是当整合来自不同来源的数据集时。这是一个非常重要的问题,广泛地称为实体对齐或匹配,并且近年来,诸如图匹配之类的问题的子集已经非常流行。为了解决这个问题,目前的工作计算对齐的基础上现有的数据嵌入。如果我们的最终目标是将两个数据集联合分析到相同的潜在因子空间中,这可能会有问题:每个数据集单独计算的嵌入可能会产生次优对齐,并且如果这样的对齐用于随后计算联合潜在因子,则计算将同样受到不完美对齐引起的复合误差的困扰。在这项工作中,我们是第一个定义和解决联合张量对齐和分解到一个共享的潜在空间的问题。通过将其作为一个统一的问题并同时解决这两个任务,我们观察到对齐和因子分解任务彼此受益,从而与两阶段方法相比具有上级性能。我们广泛地评估了我们提出的方法TENALIGN,并进行了彻底的灵敏度和消融分析。我们证明,TENALIGN显着优于基线方法,其中嵌入和匹配分别发生。
Multimodal datasets represented as tensors oftentimes share some of their modes. However, even though there may exist a one-to-one (or perhaps partial) correspondence between the coupled modes, such correspondence/alignment may not be given, especially when integrating datasets from disparate sources. This is a very important problem, broadly termed as entity alignment or matching, and subsets of the problem such as graph matching have been extremely popular in the recent years. In order to solve this problem, current work computes the alignment based on existing embeddings of the data. This can be problematic if our end goal is the joint analysis of the two datasets into the same latent factor space: the embeddings computed separately per dataset may yield a suboptimal alignment, and if such an alignment is used to subsequently compute the joint latent factors, the computation will similarly be plagued by compounding errors incurred by the imperfect alignment. In this work, we are the first to define and solve the problem of joint tensor alignment and factorization into a shared latent space. By posing this as a unified problem and solving for both tasks simultaneously, we observe that the both alignment and factorization tasks benefit each other resulting in superior performance compared to two-stage approaches. We extensively evaluate our proposed method TENALIGN and conduct a thorough sensitivity and ablation analysis. We demonstrate that TENALIGN significantly outperforms baseline approaches where embedding and matching happen separately.