We Don't Know What We Don't Know: When and How the Use of Twitter's Public APIs Biases Scientific Inference

We Don't Know What We Don't Know: When and How the Use of Twitter's Public APIs Biases Scientific Inference
复制标题

我们不知道我们不知道什么:Twitter 公共 API 的使用何时以及如何影响科学推理

DOI:
10.2139/ssrn.3079927
复制
发表时间:
2017
期刊:
PSN: Computational Models (Quantitative) (Topic)
影响因子:
--
通讯作者:
D. Stockmann
D. Stockmann
中科院分区:
--
文献类型:
--
作者:
Rebekah Tromble;A. Storz;D. Stockmann

文献摘要

被引文献

相似文献

尽管 Twitter 的研究激增,但数据收集的标准尚未具体化。使用关键字查询时,最常见的数据源(搜索和流 API)很少返回完整的推文群体,学者们不知道他们的数据是否构成代表性样本。本文旨在对可能导致的潜在偏见提供迄今为止最全面的看法。我们将来自四个相同关键字查询的数据运用到 Firehose(它提供全部推文,但成本高昂)、流媒体和搜索 API,并使用 Kendall’s-tau 和 logit 回归分析来了解数据集中的差异,包括哪些用户和内容特征使推文更有可能出现在采样结果中。我们发现,在我们检查的几乎所有数据集中,确实存在系统性差异,可能会导致学者的研究结果出现偏差,因此我们建议在未来的 Twitter 研究中要格外谨慎。
Though Twitter research has proliferated, no standards for data collection have crystallized. When using keyword queries, the most common data sources—the Search and Streaming APIs—rarely return the full population of tweets, and scholars do not know whether their data constitute a representative sample. This paper seeks to provide the most comprehensive look to-date at the potential biases that may result. Employing data derived from four identical keyword queries to the Firehose (which provides the full population of tweets but is cost-prohibitive), Streaming, and Search APIs, we use Kendall’s-tau and logit regression analyses to understand the differences in the datasets, including what user and content characteristics make a tweet more or less likely to appear in sampled results. We find that there are indeed systematic differences that are likely to bias scholars’ findings in almost all datasets we examine, and we recommend significant caution in future Twitter research.