We Don't Know What We Don't Know: When and How the Use of Twitter's Public APIs Biases Scientific Inference
We Don't Know What We Don't Know: When and How the Use of Twitter's Public APIs Biases Scientific Inference
复制标题
我们不知道我们不知道什么:Twitter 公共 API 的使用何时以及如何影响科学推理
DOI:
10.2139/ssrn.3079927
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
D. Stockmann
中科院分区:
文献类型:
--
作者:
Rebekah Tromble;A. Storz;D. Stockmann
Though Twitter research has proliferated, no standards for data collection have crystallized. When using keyword queries, the most common data sources—the Search and Streaming APIs—rarely return the full population of tweets, and scholars do not know whether their data constitute a representative sample. This paper seeks to provide the most comprehensive look to-date at the potential biases that may result. Employing data derived from four identical keyword queries to the Firehose (which provides the full population of tweets but is cost-prohibitive), Streaming, and Search APIs, we use Kendall’s-tau and logit regression analyses to understand the differences in the datasets, including what user and content characteristics make a tweet more or less likely to appear in sampled results. We find that there are indeed systematic differences that are likely to bias scholars’ findings in almost all datasets we examine, and we recommend significant caution in future Twitter research.