Graph Regularized Transductive Classification on Heterogeneous Information Networks

Graph Regularized Transductive Classification on Heterogeneous Information Networks
复制标题

DOI:
10.1007/978-3-642-15880-3_42
复制
发表时间:
2010-09
期刊:
--
影响因子:
--
通讯作者:
Ming Ji;Yizhou Sun;Marina Danilevsky;Jiawei Han;Jing Gao
Ming Ji;Yizhou Sun;Marina Danilevsky;Jiawei Han;Jing Gao
中科院分区:
其他
文献类型:
--
作者:
Ming Ji;Yizhou Sun;Marina Danilevsky;Jiawei Han;Jing Gao

文献摘要

被引文献

相似文献

异构信息网络是由多种类型的对象和链路组成的网络。最近,人们认识到强类型异构信息网络在现实世界中非常普遍。有时,某些对象可以使用标签信息。通过转导分类从这些标记和未标记的数据中学习,可以很好地提取隐藏网络结构的知识。然而,尽管同质网络上的分类已经研究了几十年,但异构网络上的分类直到最近才开始探索。本文研究了具有共同主题的异构网络数据的转换分类问题。在给定的网络中,只有一些对象被标记,我们的目标是预测所有类型的剩余对象的标签。针对具有任意网络模式和任意数量对象/链路类型的信息网络,提出了一种新的基于图的正则化框架GNetMine。具体来说,我们通过保持分别对应于每种类型链接的每个关系图的一致性,显式地尊重类型差异。然后引入有效的计算方案来解决相应的优化问题。在DBLP数据集上的实验表明,我们的算法比现有的最先进的方法显著提高了分类精度。
A heterogeneous information network is a network composed of multiple types of objects and links. Recently, it has been recognized that strongly-typed heterogeneous information networks are prevalent in the real world. Sometimes, label information is available for some objects. Learning from such labeled and unlabeled data via transductive classification can lead to good knowledge extraction of the hidden network structure. However, although classification on homogeneous networks has been studied for decades, classification on heterogeneous networks has not been explored until recently.In this paper, we consider the transductive classification problem on heterogeneous networked data which share a common topic. Only some objects in the given network are labeled, and we aim to predict labels for all types of the remaining objects. A novel graph-based regularization framework, GNetMine, is proposed to model the link structure in information networks with arbitrary network schema and arbitrary number of object/link types. Specifically, we explicitly respect the type differences by preserving consistency over each relation graph corresponding to each type of links separately. Efficient computational schemes are then introduced to solve the corresponding optimization problem. Experiments on the DBLP data set show that our algorithm significantly improves the classification accuracy over existing state-of-the-art methods.