Easing Embedding Learning by Comprehensive Transcription of Heterogeneous Information Networks

Easing Embedding Learning by Comprehensive Transcription of Heterogeneous Information Networks
复制标题

DOI:
10.1145/3219819.3220006
复制
发表时间:
2018-07
期刊:
Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Yu Shi;Qi Zhu;Fang Guo;Chao Zhang;Jiawei Han
Yu Shi;Qi Zhu;Fang Guo;Chao Zhang;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Yu Shi;Qi Zhu;Fang Guo;Chao Zhang;Jiawei Han

文献摘要

被引文献

相似文献

异构信息网络(HIN)在现实世界的应用中无处不在。与此同时,网络嵌入已经成为从网络数据中挖掘和学习的方便工具。因此,开发HIN嵌入方法是有意义的。然而,HIN中的异构性不仅引入了丰富的信息,而且还引入了潜在的不兼容语义,这对HIN中的嵌入式学习提出了特殊的挑战。为了保留HIN嵌入中丰富但潜在不兼容的信息,我们提出研究异构信息网络的综合转录问题。HIN的全面转录还提供了一种易于使用的方法来释放HIN的力量,因为它不需要额外的监督,专业知识或功能工程。为了科普在全面转录的HIN的挑战,我们提出了HEER算法,它嵌入HIN通过边缘表示,进一步耦合适当学习的异构指标。为了证实HEER的有效性,我们在两个大规模的真实词数据集上进行了实验,并进行了边缘重建任务和多个案例研究。实验结果证明了所提出的HEER模型的有效性以及边缘表示和异构度量的实用性。代码和数据可在https://github.com/GentleZhu/HEER上获得。
Heterogeneous information networks (HINs) are ubiquitous in real-world applications. In the meantime, network embedding has emerged as a convenient tool to mine and learn from networked data. As a result, it is of interest to develop HIN embedding methods. However, the heterogeneity in HINs introduces not only rich information but also potentially incompatible semantics, which poses special challenges to embedding learning in HINs. With the intention to preserve the rich yet potentially incompatible information in HIN embedding, we propose to study the problem of comprehensive transcription of heterogeneous information networks. The comprehensive transcription of HINs also provides an easy-to-use approach to unleash the power of HINs, since it requires no additional supervision, expertise, or feature engineering. To cope with the challenges in the comprehensive transcription of HINs, we propose the HEER algorithm, which embeds HINs via edge representations that are further coupled with properly-learned heterogeneous metrics. To corroborate the efficacy of HEER, we conducted experiments on two large-scale real-words datasets with an edge reconstruction task and multiple case studies. Experiment results demonstrate the effectiveness of the proposed HEER model and the utility of edge representations and heterogeneous metrics. The code and data are available at https://github.com/GentleZhu/HEER.