Predicting and Interpolating State-Level Polls Using Twitter Textual Data

Predicting and Interpolating State-Level Polls Using Twitter Textual Data
复制标题

DOI:
10.1111/ajps.12274
复制
发表时间:
2017-04-01
影响因子:
4.2
通讯作者:
Beauchamp, Nicholas
Beauchamp, Nicholas
中科院分区:
法学1区
文献类型:
--
作者:
Beauchamp, Nicholas

文献摘要

被引文献

相似文献

使用现有的调查方法,在空间或时间上密集的投票仍然既困难又昂贵。作为回应,人们越来越努力地使用社交媒体来近似各种调查指标,但这些方法中的大多数在方法上仍然存在缺陷。为了弥补这些缺陷,本文将2012年总统竞选期间的1200个州级民调与超过1亿条位于各州的政治推文结合在一起;使用一种新的线性正则化特征选择方法将民调建模为Twitter文本的函数;通过样本外测试表明,如果建模正确,基于Twitter的指标可以跟踪并在一定程度上预测民意调查,并可以扩展到未民调的州和潜在的子州地区和次日时间尺度。对最具预测性的文本特征的研究揭示了与观点转变相关的话题和事件,揭示了更普遍的关于注意和信息处理方面的党派差异的理论,并可能对实时竞选策略有所帮助。
Spatially or temporally dense polling remains both difficult and expensive using existing survey methods. In response, there have been increasing efforts to approximate various survey measures using social media, but most of these approaches remain methodologically flawed. To remedy these flaws, this article combines 1,200 state-level polls during the 2012 presidential campaign with over 100 million state-located political tweets; models the polls as a function of the Twitter text using a new linear regularization feature-selection method; and shows via out-of-sample testing that when properly modeled, the Twitter-based measures track and to some degree predict opinion polls, and can be extended to unpolled states and potentially substate regions and subday timescales. An examination of the most predictive textual features reveals the topics and events associated with opinion shifts, sheds light on more general theories of partisan difference in attention and information processing, and may be of use for real-time campaign strategy.