Dynamic Flows on Curved Space Generated by Labeled Data

Dynamic Flows on Curved Space Generated by Labeled Data
复制标题

DOI:
10.48550/arxiv.2302.00061
复制
发表时间:
2023-01
期刊:
--
影响因子:
--
通讯作者:
Xinru Hua;Truyen V. Nguyen;Tam Le;J. Blanchet;Viet Anh Nguyen
Xinru Hua;Truyen V. Nguyen;Tam Le;J. Blanchet;Viet Anh Nguyen
中科院分区:
其他
文献类型:
--
作者:
Xinru Hua;Truyen V. Nguyen;Tam Le;J. Blanchet;Viet Anh Nguyen

文献摘要

相似文献

标签数据的稀缺是许多机器学习任务的长期挑战。我们提出了我们的梯度流方法来利用现有的数据集(即源)来生成与感兴趣的数据集(即目标)接近的新样本。我们将两个数据集提升到特征-高斯流形上的概率分布空间,然后发展了一种最小化最大平均差异损失的梯度流方法。为了在弯曲的特征-高斯空间上实现分布的梯度流,我们揭示了空间的黎曼结构,并显式地计算了由最优输运度规引起的损失函数的黎曼梯度。对于实际应用,我们还提出了一种离散化的流,并提供了保证流全局收敛于最优的条件结果。我们给出了我们提出的梯度流方法在几个真实数据集上的结果,并表明我们的方法可以在迁移学习环境下提高分类模型的精度。
The scarcity of labeled data is a long-standing challenge for many machine learning tasks. We propose our gradient flow method to leverage the existing dataset (i.e., source) to generate new samples that are close to the dataset of interest (i.e., target). We lift both datasets to the space of probability distributions on the feature-Gaussian manifold, and then develop a gradient flow method that minimizes the maximum mean discrepancy loss. To perform the gradient flow of distributions on the curved feature-Gaussian space, we unravel the Riemannian structure of the space and compute explicitly the Riemannian gradient of the loss function induced by the optimal transport metric. For practical applications, we also propose a discretized flow, and provide conditional results guaranteeing the global convergence of the flow to the optimum. We illustrate the results of our proposed gradient flow method on several real-world datasets and show our method can improve the accuracy of classification models in transfer learning settings.