Comparing apples to oranges: a scalable solution with heterogeneous hashing

Comparing apples to oranges: a scalable solution with heterogeneous hashing
复制标题

DOI:
10.1145/2487575.2487668
复制
发表时间:
2013-08
期刊:
Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining
影响因子:
--
通讯作者:
Mingdong Ou;Peng Cui;Fei Wang;Jun Wang;Wenwu Zhu;Shiqiang Yang
Mingdong Ou;Peng Cui;Fei Wang;Jun Wang;Wenwu Zhu;Shiqiang Yang
中科院分区:
其他
文献类型:
--
作者:
Mingdong Ou;Peng Cui;Fei Wang;Jun Wang;Wenwu Zhu;Shiqiang Yang

文献摘要

被引文献

相似文献

尽管散列技术已经流行于大规模相似性搜索问题,但是用于设计最优散列函数的大多数现有方法集中于同质相似性评估,即,要被索引的数据实体是相同类型的。认识到异构实体和关系在真实的世界应用中也是普遍存在的,存在对从多个异构域(例如,向某个Facebook用户推荐相关的帖子和图像。在本文中,我们解决了大规模设置下的“比较苹果和橘子”的问题。具体来说,我们提出了一种新的异构哈希(RaHH),它提供了一个通用的框架,用于生成哈希码的数据实体坐在多个异构域。与现有的一些将异构数据映射到公共汉明空间的哈希方法不同,RaHH方法为每种类型的数据实体构建一个汉明空间,并同时学习它们之间的最佳映射。这使得学习的哈希码灵活地科普不同数据域的特性。此外,RaHH框架对数据实体之间的同构和异构关系进行编码,以设计具有更高准确性的散列函数。为了验证所提出的RaHH方法,我们对两个大型数据集进行了广泛的评估;一个是从流行的社交媒体网站,腾讯微博,另一个是Flickr(NUS-WIDE)的开放数据集。实验结果清楚地表明,RaHH优于几个国家的最先进的散列方法具有显着的性能增益。
Although hashing techniques have been popular for the large scale similarity search problem, most of the existing methods for designing optimal hash functions focus on homogeneous similarity assessment, i.e., the data entities to be indexed are of the same type. Realizing that heterogeneous entities and relationships are also ubiquitous in the real world applications, there is an emerging need to retrieve and search similar or relevant data entities from multiple heterogeneous domains, e.g., recommending relevant posts and images to a certain Facebook user. In this paper, we address the problem of ``comparing apples to oranges'' under the large scale setting. Specifically, we propose a novel Relation-aware Heterogeneous Hashing (RaHH), which provides a general framework for generating hash codes of data entities sitting in multiple heterogeneous domains. Unlike some existing hashing methods that map heterogeneous data in a common Hamming space, the RaHH approach constructs a Hamming space for each type of data entities, and learns optimal mappings between them simultaneously. This makes the learned hash codes flexibly cope with the characteristics of different data domains. Moreover, the RaHH framework encodes both homogeneous and heterogeneous relationships between the data entities to design hash functions with improved accuracy. To validate the proposed RaHH method, we conduct extensive evaluations on two large datasets; one is crawled from a popular social media sites, Tencent Weibo, and the other is an open dataset of Flickr(NUS-WIDE). The experimental results clearly demonstrate that the RaHH outperforms several state-of-the-art hashing methods with significant performance gains.