Diffusion Transport Alignment

Diffusion Transport Alignment
复制标题

DOI:
10.48550/arxiv.2206.07305
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Andres F. Duque;Guy Wolf;Kevin R. Moon
Andres F. Duque;Guy Wolf;Kevin R. Moon
中科院分区:
其他
文献类型:
--
作者:
Andres F. Duque;Guy Wolf;Kevin R. Moon

文献摘要

相似文献

当用不同的工具或条件对某一现象进行研究产生不同但相关的领域时,多模态数据的整合就提出了挑战。许多现有的数据集成方法假设整个数据集的域之间存在已知的一一对应关系,这可能是不现实的。此外,现有的流形对齐方法不适合于数据包含域特定区域的情况,即,在另一个域中不存在用于数据的某个部分的对应物。我们提出了扩散传输对齐(DTA),一种半监督流形对齐方法,利用只有几个点之间的先验对应知识来对齐域。通过建立一个扩散过程,DTA找到一个传输计划之间的测量数据从两个异构域具有不同的特征空间,通过假设,共享一个类似的几何结构来自相同的底层数据生成过程。DTA还可以以数据驱动的方式计算部分比对,从而在仅在一个域中测量某些数据时产生准确的比对。我们经验证明,DTA在半监督环境中对齐多模态数据方面优于其他方法。我们还通过经验表明,DTA获得的对齐可以提高机器学习任务的性能,例如域自适应,域间特征映射和探索性数据分析,同时优于竞争方法。
The integration of multimodal data presents a challenge in cases when the study of a given phenomena by different instruments or conditions generates distinct but related domains. Many existing data integration methods assume a known one-to-one correspondence between domains of the entire dataset, which may be unrealistic. Furthermore, existing manifold alignment methods are not suited for cases where the data contains domain-specific regions, i.e., there is not a counterpart for a certain portion of the data in the other domain. We propose Diffusion Transport Alignment (DTA), a semi-supervised manifold alignment method that exploits prior correspondence knowledge between only a few points to align the domains. By building a diffusion process, DTA finds a transportation plan between data measured from two heterogeneous domains with different feature spaces, which by assumption, share a similar geometrical structure coming from the same underlying data generating process. DTA can also compute a partial alignment in a data-driven fashion, resulting in accurate alignments when some data are measured in only one domain. We empirically demonstrate that DTA outperforms other methods in aligning multimodal data in this semisupervised setting. We also empirically show that the alignment obtained by DTA can improve the performance of machine learning tasks, such as domain adaptation, inter-domain feature mapping, and exploratory data analysis, while outperforming competing methods.