Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment

Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment
复制标题

DOI:
10.1177/1948550619876629
复制
发表时间:
2020-02-19
影响因子:
5.7
通讯作者:
Dehghani, Morteza
Dehghani, Morteza
中科院分区:
心理学2区
文献类型:
--
作者:
Hoover, Joe;Portillo-Wightman, Gwenyth;Dehghani, Morteza

文献摘要

被引文献

相似文献

研究表明,在自然语言中解释道德情感可以洞察各种在线和离线现象,如消息传播,抗议动态和社交距离。然而,用自然语言测量道德情感是一项挑战,而且由于注释数据的有限性,这项任务的难度更大。为了解决这个问题,我们介绍了道德基础Twitter语料库,收集了35,108条推文,这些推文来自七个不同的话语领域,并由至少三名训练有素的注释者为10类道德情感进行了手工注释。为了方便调查的注释者的反应动力学,我们还提供了每个注释者的心理和人口统计元数据。最后,我们使用一系列流行的方法报告道德情感分类基线的语料库。
Research has shown that accounting for moral sentiment in natural language can yield insight into a variety of on- and off-line phenomena such as message diffusion, protest dynamics, and social distancing. However, measuring moral sentiment in natural language is challenging, and the difficulty of this task is exacerbated by the limited availability of annotated data. To address this issue, we introduce the Moral Foundations Twitter Corpus, a collection of 35,108 tweets that have been curated from seven distinct domains of discourse and hand annotated by at least three trained annotators for 10 categories of moral sentiment. To facilitate investigations of annotator response dynamics, we also provide psychological and demographic metadata for each annotator. Finally, we report moral sentiment classification baselines for this corpus using a range of popular methodologies.