Supervised Tree-Wasserstein Distance

Supervised Tree-Wasserstein Distance
复制标题

DOI:
--
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuki Takezawa;R. Sato;M. Yamada
Yuki Takezawa;R. Sato;M. Yamada
中科院分区:
其他
文献类型:
--
作者:
Yuki Takezawa;R. Sato;M. Yamada

文献摘要

相似文献

为了衡量文档的相似度,Wasserstein距离是一个强大的工具,但它需要很高的计算成本。最近,为了快速计算 Wasserstein 距离,已经提出了使用树度量来近似 Wasserstein 距离的方法。这些基于树的方法可以快速比较大量文档;然而,它们不受监督,也不会学习特定任务的距离。在这项工作中,我们提出了监督树-Wasserstein(STW)距离,这是一种基于树度量的快速、监督度量学习方法。具体来说,我们通过树的父子关系重写树度量上的 Wasserstein 距离,并使用对比损失将其表示为连续优化问题。实验表明,STW 距离可以快速计算,并提高了文档分类任务的准确性。此外,STW距离是通过矩阵乘法制定的,在GPU上运行,并且适合批处理。因此,我们表明 STW 距离在比较大量文档时非常有效。
To measure the similarity of documents, the Wasserstein distance is a powerful tool, but it requires a high computational cost. Recently, for fast computation of the Wasserstein distance, methods for approximating the Wasserstein distance using a tree metric have been proposed. These tree-based methods allow fast comparisons of a large number of documents; however, they are unsupervised and do not learn task-specific distances. In this work, we propose the Supervised Tree-Wasserstein (STW) distance, a fast, supervised metric learning method based on the tree metric. Specifically, we rewrite the Wasserstein distance on the tree metric by the parent-child relationships of a tree and formulate it as a continuous optimization problem using a contrastive loss. Experimentally, we show that the STW distance can be computed fast, and improves the accuracy of document classification tasks. Furthermore, the STW distance is formulated by matrix multiplications, runs on a GPU, and is suitable for batch processing. Therefore, we show that the STW distance is extremely efficient when comparing a large number of documents.