Claim Verification under Positive Unlabeled Learning

Claim Verification under Positive Unlabeled Learning
复制标题

DOI:
10.1109/asonam49781.2020.9381336
复制
发表时间:
2020-12
期刊:
2020 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)
影响因子:
--
通讯作者:
Fan Yang;E. Dragut;Arjun Mukherjee
Fan Yang;E. Dragut;Arjun Mukherjee
中科院分区:
其他
文献类型:
--
作者:
Fan Yang;E. Dragut;Arjun Mukherjee

文献摘要

被引文献

相似文献

我们将证据感知声明验证扩展到正-未标记(PU)学习的上下文中。已有的工作假设索赔的真假是已知的,并将其作为一个有监督的学习问题来形成任务。然而,这一假设低估了收集虚假索赔的难度;我们认为,在没有负面标签的情况下,索赔验证更具挑战性。我们考虑一种更实际的设置,在这种情况下,只有相对较少的真实声明被标记,而更多的声明仍未标记。因此,我们将索赔验证任务描述为PU学习问题。我们将声明-证据对的学习表示从PU学习中分离出来,并采用预先训练的通用语言模型对声明-证据对进行编码。我们进一步提出使用产生式对抗网络(GAN)来捕捉编码声明-证据对与真实性之间的潜在对齐。我们通过扩展以前基于GaN的PU学习来利用验证作为GAN的一部分。结果表明,该模型在使用少量标记数据的情况下取得了最好的性能,并且对真实性先验估计具有较强的稳健性。我们对模型的选择进行了深入的分析。该方法在两种实际情况下表现最好:(I)未标记数据多于已标记数据;(Ii)未标记正数据多于未标记负数据。
We extend evidence-aware claim verification to the context of positive-unlabeled (PU) learning. Existing works assume the truth and the falsity of the claims are known for training and form the task as a supervised learning problem. However, this assumption underestimates the difficulty of collecting false claims; we argue that claim verification is more challenging in the absence of negative labels. We consider a more practical setting, where only a comparatively small number of true claims are labeled and more claims remain unlabeled. Thus, we formulate the claim verification task as a PU learning problem. We decouple learning representation of claim-evidence pair from PU learning and adopt a pre-trained universal language model to encode claim-evidence pairs. We further propose to use the generative adversarial network (GAN) to capture the latent alignment between encoded claim-evidence pair and the truthfulness. We leverage the verification as part of the GAN by extending previous GAN based PU learning. We show that the proposed model achieves the best performance with a small amount of labeled data and is robust to the truthfulness prior estimation. We conduct a thorough analysis of the model selection. The proposed approach performs the best under two practical scenarios: (i) the unlabeled data is more than the labeled data; (ii) and the unlabeled positive data is more than the unlabeled negative data.