TIMME: Twitter Ideology-detection via Multi-task Multi-relational Embedding

TIMME: Twitter Ideology-detection via Multi-task Multi-relational Embedding
复制标题

DOI:
10.1145/3394486.3403275
复制
发表时间:
2020-06
期刊:
Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Zhiping Xiao;Weiping Song;Haoyan Xu;Zhicheng Ren;Yizhou Sun
Zhiping Xiao;Weiping Song;Haoyan Xu;Zhicheng Ren;Yizhou Sun
中科院分区:
其他
文献类型:
--
作者:
Zhiping Xiao;Weiping Song;Haoyan Xu;Zhicheng Ren;Yizhou Sun

文献摘要

被引文献

相似文献

我们的目标是解决预测人们的意识形态或政治倾向的问题。我们使用Twitter数据对其进行估计,并将其形式化为分类问题。意识形态侦查一直是一个既具有挑战性又十分重要的问题。某些群体,如政策制定者,依靠它来做出明智的决定。在过去,收集民意需要大量的调查研究,分析普通公民的政治倾向是不容易的。Twitter等社交媒体的兴起,使我们能够轻松地收集普通公民的数据。然而,社交网络数据集中标签和特征的不完全性很棘手,更不用说庞大的数据规模和异构性了。这些数据与许多常用的数据集有很大的不同,因此带来了独特的挑战。在我们的工作中,首先我们从Twitter建立了我们自己的数据集。其次,我们提出了一种多任务多关系嵌入模型TIMME,该模型能够有效地处理稀疏标记的异构现实数据集。它还可以处理输入特征的不完整性。实验结果表明,TIMME总体上优于最先进的Twitter意识形态检测模型。我们的发现包括:链接可以在没有文本的情况下产生良好的分类结果;保守的声音在推特上没有被充分代表;跟随是预测意识形态最重要的关系;转发和提及可以提高点赞的机会,等等。最后但并非最不重要的一点是,理论上TIMME可以扩展到其他数据集和任务。
We aim at solving the problem of predicting people's ideology, or political tendency. We estimate it by using Twitter data, and formalize it as a classification problem. Ideology-detection has long been a challenging yet important problem. Certain groups, such as the policy makers, rely on it to make wise decisions. Back in the old days when labor-intensive survey-studies were needed to collect public opinions, analyzing ordinary citizens' political tendencies was uneasy. The rise of social medias, such as Twitter, has enabled us to gather ordinary citizen's data easily. However, the incompleteness of the labels and the features in social network datasets is tricky, not to mention the enormous data size and the heterogeneousity. The data differ dramatically from many commonly-used datasets, thus brings unique challenges. In our work, first we built our own datasets from Twitter. Next, we proposed TIMME, a multi-task multi-relational embedding model, that works efficiently on sparsely-labeled heterogeneous real-world dataset. It could also handle the incompleteness of the input features. Experimental results showed that TIMME is overall better than the state-of-the-art models for ideology detection on Twitter. Our findings include: links can lead to good classification outcomes without text; conservative voice is under-represented on Twitter; follow is the most important relation to predict ideology; retweet and mention enhance a higher chance of like, etc. Last but not least, TIMME could be extended to other datasets and tasks in theory.