Query-dependent cross-domain ranking in heterogeneous network

Query-dependent cross-domain ranking in heterogeneous network
复制标题

DOI:
10.1007/s10115-011-0472-7
复制
发表时间:
2012-01
影响因子:
2.7
通讯作者:
Bo Wang;Jie Tang;Wei Fan;Songcan Chen;Chenhao Tan;Zi Yang
Bo Wang;Jie Tang;Wei Fan;Songcan Chen;Chenhao Tan;Zi Yang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Bo Wang;Jie Tang;Wei Fan;Songcan Chen;Chenhao Tan;Zi Yang

文献摘要

相似文献

传统的学习排序问题主要集中在单一类型的对象上。然而,随着Web 2.0的快速增长,在多个相互关联的异质对象上进行排名成为一种常见的情况,例如,异质学术网络。在这种情况下,对于某些类型的对象(例如,会议)可能有很多训练数据,而对于感兴趣的对象类型(例如,作者)只有很少的训练数据。因此,两个重要的问题是:(1)给定一个网络化的数据集,如何借鉴其他类型的对象的监督,以便在监督不足的情况下为感兴趣的对象建立准确的排序模型?(2)如果不同对象之间存在链接,如何利用它们之间的关系来提高排序性能?在这项工作中,我们首先提出了一个称为HCDRank的正则化框架,以同时最小化与这两个域相关的两个损失函数。然后,通过利用异类对象之间的链接信息对该方法进行了扩展。我们对所提出的方法进行了理论分析,并推导出它的广义界,以说明这两个相关领域如何在学习排名函数时相互帮助。在三种不同类型的数据集上的实验结果证明了所提方法的有效性。
Traditional learning-to-rank problem mainly focuses on one single type of objects. However, with the rapid growth of the Web 2.0, ranking over multiple interrelated and heterogeneous objects becomes a common situation, e.g., the heterogeneous academic network. In this scenario, one may have much training data for some type of objects (e.g. conferences) while only very few for the interested types of objects (e.g. authors). Thus, the two important questions are: (1) Given a networked data set, how could one borrow supervision from other types of objects in order to build an accurate ranking model for the interested objects with insufficient supervision? (2) If there are links between different objects, how can we exploit their relationships for improved ranking performance? In this work, we first propose a regularized framework called HCDRank to simultaneously minimize two loss functions related to these two domains. Then, we extend the approach by exploiting the link information between heterogeneous objects. We conduct a theoretical analysis to the proposed approach and derive its generalization bound to demonstrate how the two related domains could help each other in learning ranking functions. Experimental results on three different genres of data sets demonstrate the effectiveness of the proposed approaches.