Infer-AVAE: An attribute inference model based on adversarial variational autoencoder
Infer-AVAE: An attribute inference model based on adversarial variational autoencoder
复制标题
Infer-AVAE:基于对抗性变分自动编码器的属性推理模型
DOI:
10.1016/j.neucom.2022.02.006
复制
发表时间:
2022
期刊:
影响因子:
6
通讯作者:
Xiaohong Guan
中科院分区:
文献类型:
--
作者:
Yadong Zhou;Zhihao Ding;Xiaoming Liu;Chao Shen;Lingling Tong;Xiaohong Guan
User attributes, such as gender and education, face severe incompleteness in social networks. Attribute inference aims to infer users’ missing attribute labels based on observed data to make this valuable data usable for downstream tasks like user profiling and personalized recommendation. Recently, variational autoencoder (VAE), an end-to-end deep generative model, has shown promising performance by handling the problem in a semi-supervised way. However, VAEs can easily suffer from over-fitting and over-smoothing when applied to attribute inference. Specifically, VAE implemented with multi-layer perceptron (MLP) can only reconstruct input data but fail to infer missing parts. While using the trending graph neural networks (GNNs) as encoder has the problem that GNNs aggregate redundant information from the neighborhood and generate indistinguishable user representations, known as over-smoothing. In this paper, we propose an attribute Inference model based on Adversarial VAE (Infer-AVAE) to cope with these issues. Specifically, to overcome over-smoothing, Infer-AVAE unifies MLP and GNNs in the encoder to learn positive and negative latent representations respectively. Meanwhile, an adversarial network is trained to distinguish the two representations, and GNNs are trained to aggregate less noise for more robust representations through adversarial training. Finally, to relieve over-fitting, mutual information constraint is introduced as a regularizer for the decoder to make better use of auxiliary information in representations and generate outputs not limited by observations. We evaluate our model on four real-world social network datasets, and experimental results demonstrate that our model averagely outperforms baselines by 7.0% in accuracy.