Multi-Source Domain Adaptation with Weak Supervision for Early Fake News Detection

Multi-Source Domain Adaptation with Weak Supervision for Early Fake News Detection
复制标题

弱监督的多源域适应用于早期假新闻检测

DOI:
10.1109/bigdata52589.2021.9671592
复制
发表时间:
2021
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Kai Shu
Kai Shu
中科院分区:
--
文献类型:
--
作者:
Yichuan Li;Kyumin Lee;Nima Kordzadeh;Brenton D. Faber;Cameron Fiddes;Elaine Chen;Kai Shu

文献摘要

参考文献

被引文献

相似文献

最近,从政治到娱乐和健康的大量和多样化的假新闻放大了社会不信任问题,成为社会和研究界的一大挑战。现有的假新闻检测方法大多是针对特定领域设计的,或者需要来自各个领域的大量标记数据。如果在某个领域没有足够的标记数据,现有的模型可能无法很好地检测来自该领域的假新闻。为了克服这些限制,我们提出了一种基于多频域自适应和弱监督的早期假新闻检测新框架。该框架通过多源域自适应将已有标记的源域知识转移到具有有限甚至没有标记数据的目标/新域,并通过弱监督将研究者关于假新闻的先验知识应用到目标域。弱监督通过已知的启发式规则为目标域中的未标记样本分配弱标记。我们的实验结果表明,我们的方法优于7个国家的最先进的方法在三个真实世界的数据集。特别是,我们的模型平均比最佳基线高出5.2%的准确度。我们的模型采用了更先进的编码器,可以进一步提高3.7%的性能。代码可在此点击链接。
Recently, the massive and diverse fake news from politics to entertainment and health has amplified the social distrust problem and has become a big challenge for the society and research community. The existing fake news detection methods are mostly designed for either a specific domain or require huge labeled data from various domains. If there is not enough labeled data in a certain domain, existing models may not work well for detecting fake news from that domain. To overcome these limitations we propose a novel framework based on multisource domain adaptation and weak supervision for early fake news detection. The framework transfers sufficient labeled source domains’ knowledge into a target/new domain with limited or even no labeled data by the multi-source domain adaptation, and applies researchers’ prior knowledge about fake news to the target domain by the weak supervision. The weak supervision assigns the weak labels to the unlabeled samples in the target domain through known heuristic rules. Our experimental results show that our approach outperforms 7 state-of-the-art methods in three real-world datasets. In particular, our model achieves, on average, 5.2% higher accuracy than the best baseline. Our model with a more advanced encoder can further boost the performance by 3.7%. The code is available at this clickable link.
DOI: 10.1007/978-3-030-58545-7_33
发表时间: 2020-07
期刊: --
影响因子: --
作者:
S. Paul;Yi-Hsuan Tsai;S. Schulter;A. Roy-Chowdhury;Manmohan Chandraker
通讯作者: S. Paul;Yi-Hsuan Tsai;S. Schulter;A. Roy-Chowdhury;Manmohan Chandraker
DOI: 10.1109/bigdata47090.2019.9005556
发表时间: 2019-12
期刊: 2019 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Jiawei Zhang;Bowen Dong;Philip S. Yu
通讯作者: Jiawei Zhang;Bowen Dong;Philip S. Yu
DOI: 10.1609/aaai.v34i01.5389
发表时间: 2019-12
期刊: ArXiv
影响因子: --
作者:
Yaqing Wang;Weifeng Yang;Fenglong Ma;Jin Xu;Bin Zhong;Qiang Deng;Jing Gao
通讯作者: Yaqing Wang;Weifeng Yang;Fenglong Ma;Jin Xu;Bin Zhong;Qiang Deng;Jing Gao