Exploring Generalizability of Fine-Tuned Models for Fake News Detection

Exploring Generalizability of Fine-Tuned Models for Fake News Detection
复制标题

DOI:
10.1109/cic56439.2022.00022
复制
发表时间:
2022-12
期刊:
2022 IEEE 8th International Conference on Collaboration and Internet Computing (CIC)
影响因子:
--
通讯作者:
Abhijit Suprem;Sanjyot Vaidya;C. Pu
Abhijit Suprem;Sanjyot Vaidya;C. Pu
中科院分区:
其他
文献类型:
--
作者:
Abhijit Suprem;Sanjyot Vaidya;C. Pu

文献摘要

被引文献

相似文献

新型冠状病毒肺炎(COVID-19,即2019冠状病毒病)大流行导致危险的错误信息急剧增加,CDC和世卫组织称之为“信息流行病”。与新型冠状病毒疫情相关的错误信息不断变化;这可能导致概念漂移导致微调模型的性能下降。如果模型能够很好地概括漂移数据的某些周期性方面,则可以减轻退化。在本文中,我们探索了9个假新闻数据集上预训练和微调的假新闻检测器的通用性。我们发现,现有的模型通常在训练数据集上过拟合,并且在看不见的数据上表现不佳。然而,在一些与训练数据重叠的不可见数据子集上,模型具有更高的准确性。基于这一观察,我们还提出了KMeans-Proxy,这是一种基于K-Means聚类的快速有效的方法,用于快速识别这些重叠的未知数据子集。KMeans-Proxy将不可见的假新闻数据集的泛化能力提高了0.1-0.2个f1点。我们提出了我们的泛化实验以及KMeans-Proxy,以进一步研究解决假新闻问题。
The Covid-19 pandemic has caused a dramatic and parallel rise in dangerous misinformation, denoted an ‘infodemic’ by the CDC and WHO. Misinformation tied to the Covid-19 infodemic changes continuously; this can lead to performance degradation of fine-tuned models due to concept drift. Degredation can be mitigated if models generalize well-enough to capture some cyclical aspects of drifted data. In this paper, we explore generalizability of pre-trained and fine-tuned fake news detectors across 9 fake news datasets. We show that existing models often overfit on their training dataset and have poor performance on unseen data. However, on some subsets of unseen data that overlap with training data, models have higher accuracy. Based on this observation, we also present KMeans-Proxy, a fast and effective method based on K-Means clustering for quickly identifying these overlapping subsets of unseen data. KMeans-Proxy improves generalizability on unseen fake news datasets by 0.1-0.2 f1-points across datasets. We present both our generalizability experiments as well as KMeans-Proxy to further research in tackling the fake news problem.