CAVES: A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines

CAVES: A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines
复制标题

CAVES:一个有助于对新冠疫苗问题进行可解释分类和总结的数据集

DOI:
10.1145/3477495.3531745
复制
发表时间:
2022
期刊:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Saptarshi Ghosh
Saptarshi Ghosh
中科院分区:
--
文献类型:
--
作者:
Soham Poddar;Azlaan Mustafa Samad;Rajdeep Mukherjee;Niloy Ganguly;Saptarshi Ghosh

文献摘要

参考文献

被引文献

相似文献

说服人们接种COVID-19疫苗是当今社会的一项关键挑战。作为实现这一目标的第一步,许多先前的工作依赖于社交媒体分析来了解人们对这些疫苗的具体关注,例如潜在的副作用,无效性,政治因素等。尽管有数据集将社交媒体帖子大致分为Anti-Vax和Pro-Vax标签,(据我们所知)没有数据集根据帖子中提到的具体反疫苗问题来标记社交媒体帖子。在本文中,我们策划了CAVES,这是第一个大规模数据集,包含约10 k条COVID-19反疫苗推文,在多标签环境中标记为各种特定的反疫苗问题。这也是第一个为每个标签提供解释的多标签分类数据集。此外,该数据集还提供了所有推文的类汇总。我们还对数据集进行了初步实验,并表明这对于多标签可解释分类和推文摘要来说是一个非常具有挑战性的数据集,正如一些最先进的模型所获得的中等分数所证明的那样。
Convincing people to get vaccinated against COVID-19 is a key societal challenge in the present times. As a first step towards this goal, many prior works have relied on social media analysis to understand the specific concerns that people have towards these vaccines, such as potential side-effects, ineffectiveness, political factors, and so on. Though there are datasets that broadly classify social media posts into Anti-vax and Pro-Vax labels, there is no dataset (to our knowledge) that labels social media posts according to the specific anti-vaccine concerns mentioned in the posts. In this paper, we have curated CAVES, the first large-scale dataset containing about 10k COVID-19 anti-vaccine tweets labelled into various specific anti-vaccine concerns in a multi-label setting. This is also the first multi-label classification dataset that provides explanations for each of the labels. Additionally, the dataset also provides class-wise summaries of all the tweets. We also perform preliminary experiments on the dataset and show that this is a very challenging dataset for multi-label explainable classification and tweet summarization, as is evident by the moderate scores achieved by some state-of-the-art models.
DOI: 10.18653/v1/2020.acl-main.408
发表时间: 2019-11
期刊: --
影响因子: --
作者:
Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace
通讯作者: Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace