Estimating geographic subjective well-being from Twitter: A comparison of dictionary and data-driven language methods

Estimating geographic subjective well-being from Twitter: A comparison of dictionary and data-driven language methods
复制标题

DOI:
10.1073/pnas.1906364117
复制
发表时间:
2020-05-12
影响因子:
11.1
通讯作者:
Eichstaedt, Johannes C.
Eichstaedt, Johannes C.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Jaidka, Kokil;Giorgi, Salvatore;Eichstaedt, Johannes C.

文献摘要

被引文献

相似文献

世界各地的研究人员和政策制定者都对衡量人口的主观幸福感感兴趣。当用户在社交媒体上发帖时,他们会留下反映他们想法和感受的数字痕迹。这些数字痕迹的聚合可能使大规模监测健康状况成为可能。然而,如果要产生可靠的估计,基于社交媒体的方法需要对区域影响具有稳健性。使用15.3亿条地理标记的英语推文样本,我们对单词级和数据驱动的文本分析方法进行了系统评估,以生成1208个美国县的福祉估计。我们将基于推特的县级估计与盖洛普- share - care幸福指数调查提供的173万次电话调查结果进行了比较。我们发现,由于语言使用的地区、文化和社会经济差异,单词级方法(例如,语言学调查和单词计数[LIWC] 2015和Mechanical Turk [LabMT]的语言评估)产生了不一致的县级幸福感测量。然而,只要去掉三个最常见的单词,就能显著提高幸福感预测。数据驱动的方法提供了可靠的估计,接近盖洛普数据的r = 0.64。我们表明,这些发现推广到县的社会经济和健康结果,并且在对样本进行后分层以更能代表美国一般人群时是稳健的。当使用监督数据驱动的方法时,来自社交媒体数据的区域福祉估计似乎是稳健的。
Researchers and policy makers worldwide are interested in measuring the subjective well-being of populations. When users post on social media, they leave behind digital traces that reflect their thoughts and feelings. Aggregation of such digital traces may make it possible to monitor well-being at large scale. However, social media-based methods need to be robust to regional effects if they are to produce reliable estimates. Using a sample of 1.53 billion geotagged English tweets, we provide a systematic evaluation of word-level and data-driven methods for text analysis for generating well-being estimates for 1,208 US counties. We compared Twitter-based county-level estimates with well-being measurements provided by the Gallup-Sharecare Well-Being Index survey through 1.73 million phone surveys. We find that word-level methods (e.g., Linguistic Inquiry and Word Count [LIWC] 2015 and Language Assessment by Mechanical Turk [LabMT]) yielded inconsistent county-level wellbeing measurements due to regional, cultural, and socioeconomic differences in language use. However, removing as few as three of the most frequent words led to notable improvements in well-being prediction. Data-driven methods provided robust estimates, approximating the Gallup data at up to r = 0.64. We show that the findings generalized to county socioeconomic and health outcomes and were robust when poststratifying the samples to be more representative of the general US population. Regional well-being estimation from social media data seems to be robust when supervised data-driven methods are used.