Understanding demographic and socioeconomic biases of geotagged Twitter users at the county level

Understanding demographic and socioeconomic biases of geotagged Twitter users at the county level
复制标题

DOI:
10.1080/15230406.2018.1434834
复制
发表时间:
2019-05-04
影响因子:
2.5
通讯作者:
Ye, Xinyue
Ye, Xinyue
中科院分区:
地球科学3区
文献类型:
--
作者:
Jiang, Yuqin;Li, Zhenlong;Ye, Xinyue

文献摘要

被引文献

相似文献

微博平台产生的海量社交媒体数据为研究前所未有的规模的人类动力学提供了新的数据源。与此同时,带有地理标签的 Twitter 用户的人口偏见已得到广泛认可。了解 Twitter 用户的人口统计和社会经济偏见对于对人们的态度和行为做出可靠的推断至关重要。然而,现有的全球模型无法捕捉人口和社会经济偏差的区域差异。为了弥合这一差距,我们对整个美国本土的不同人口/社会经济因素与地理标记 Twitter 用户之间的关系进行了建模,旨在了解人口和社会经济因素与县级 Twitter 用户数量之间的关系。为了有效地识别美国每个县的本地 Twitter 用户,我们集成了三种常用的方法,并在高性能计算环境中开发了一种查询方法。结果表明,我们不仅可以确定人口和社会经济因素与 Twitter 用户数量的关系,还可以测量和绘制这些因素的影响在不同县之间的差异。
Massive social media data produced from microblog platforms provide a new data source for studying human dynamics at an unprecedented scale. Meanwhile, population bias in geotagged Twitter users is widely recognized. Understanding the demographic and socioeconomic biases of Twitter users is critical for making reliable inferences on the attitudes and behaviors of the population. However, the existing global models cannot capture the regional variations of the demographic and socioeconomic biases. To bridge the gap, we modeled the relationships between different demographic/socioeconomic factors and geotagged Twitter users for the whole contiguous United States, aiming to understand how the demographic and socioeconomic factors relate to the number of Twitter users at county level. To effectively identify the local Twitter users for each county of the United States, we integrate three commonly used methods and develop a query approach in a high-performance computing environment. The results demonstrate that we can not only identify how the demographic and socioeconomic factors relate to the number of Twitter users, but can also measure and map how the influence of these factors vary across counties.