Distinguishing the Wood from the Trees: Contrasting Collection Methods to Understand Bias in a Longitudinal Brexit Twitter Dataset

Distinguishing the Wood from the Trees: Contrasting Collection Methods to Understand Bias in a Longitudinal Brexit Twitter Dataset
复制标题

区分木材和树木:对比收集方法以了解纵向英国脱欧 Twitter 数据集中的偏差

DOI:
10.1609/icwsm.v11i1.14924
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Laura Cram
Laura Cram
中科院分区:
--
文献类型:
--
作者:
Claire Llewellyn;Laura Cram

文献摘要

被引文献

相似文献

可以使用各种方法来搜索或流式传输Twitter数据,以收集特定主题的样本。所有这些方法都会在所得数据集中引入偏差。在这里,我们检查,并试图定义,不同的策略引入的偏见。理解偏差意味着我们可以以更精确的方式从数据中推断出更广泛的含义。我们使用了2016年英国-欧盟脱欧公投主题收集的数据集。每个数据集都从Twitter上提取了从2015年9月1日到2016年8月31日的12个月期间的数据。三种数据收集策略被认为是:收集人类定义的主题特定的标签;收集使用半自动化技术,以确定主题术语,然后用于收集推文;和收集从预定义的用户已知的推文的主题。为了调查数据中的偏差,我们查看并发现以下方面的广泛差异:组级元数据属性,如数据集的大小;每组中的用户数量;朋友和追随者的平均数量;可能的转发状态;以及各种附加组件的包含程度,如标签,URL和媒体。我们还发现,主题的相关性在不同的集合之间存在差异;在已知用户集合中要高得多。我们调查如何可读性的推文在每个集合的变化,特别是已知的用户和主题术语集之间。我们还发现,有一个令人惊讶的缺乏重叠的数据使用不同的收集方法。
Various methods can be used for searching or streaming Twitter data to gather a sample on a specific topic. All of these methods introduce a bias into the resulting datasets. Here we examine, and try to define, the bias that the different strategies introduce. Understanding the bias means that we can extrapolate wider meaning from the data in a more precise manner. We use datasets collected on topics from the UK-EU Brexit referendum conducted in 2016. Each dataset discussed draws data from Twitter over a twelve-month period, from 1st September 2015 until 31st August 2016. Three data collection strategies are considered: collecting on human defined topic specific hashtags; collecting using a semi-automated technique to identify topic terms which are then used to collect tweets; and collecting from predefined users known to be tweeting on the topic. To investigate bias in the data we look at, and find wide variation in: group level metadata attributes such as size of the dataset; number of users in each set; average numbers of friends and followers; likely re-tweet status; and levels of inclusion of various add-ons such as hashtags, URLs and media. We also find that relevance to the topic differs between the sets; being far higher in the known users set. We investigate how readability of tweets within each set varies, particularly between known users and topic term sets. We also find that there is a surprising lack of overlap in the data obtained using different collection methods.