What You Can Scrape and What Is Right to Scrape: A Proposal for a Tool to Collect Public Facebook Data

What You Can Scrape and What Is Right to Scrape: A Proposal for a Tool to Collect Public Facebook Data
复制标题

DOI:
10.1177/2056305120940703
复制
发表时间:
2020-07-01
影响因子:
5.2
通讯作者:
Vegetti, Federico
Vegetti, Federico
中科院分区:
人文科学2区
文献类型:
--
作者:
Mancosu, Moreno;Vegetti, Federico

文献摘要

被引文献

相似文献

作为对剑桥分析公司丑闻的回应,Facebook限制了对其应用程序编程接口(API)的访问。这项新政策破坏了独立研究人员研究政治和社会行为相关主题的可能性。然而,研究人员可能感兴趣的许多公共信息仍然可以在Facebook上获得,并且仍然可以通过网络抓取技术系统地收集。本文的目的有两个。首先,我们讨论了研究人员在计划收集和可能发布Facebook数据时应该考虑的一些伦理和法律问题。特别是,我们讨论了什么样的信息可以合乎道德地收集用户(公共信息),发布的数据应该如何遵守隐私法规(如GDPR),以及违反Facebook服务条款可能给研究人员带来的后果。其次,我们提出了一个公开的Facebook帖子的抓取程序,并讨论了一些可以执行的技术调整,以使数据在道德和法律上可接受。该代码使用屏幕抓取来收集对Facebook公开帖子的反应列表,并对用户的标识符执行单向加密散列函数,以假名化他们的个人信息,同时仍然保持数据中的可追踪性。这篇文章有助于围绕互联网研究自由的辩论,以及从社交网络上抓取数据可能引起的伦理问题。
In reaction to the Cambridge Analytica scandal, Facebook has restricted the access to its Application Programming Interface (API). This new policy has damaged the possibility for independent researchers to study relevant topics in political and social behavior. Yet, much of the public information that the researchers may be interested in is still available on Facebook, and can be still systematically collected through web scraping techniques. The goal of this article is twofold. First, we discuss some ethical and legal issues that researchers should consider as they plan their collection and possible publication of Facebook data. In particular, we discuss what kind of information can be ethically gathered about the users (public information), how published data should look like to comply with privacy regulations (like the GDPR), and what consequences violating Facebook's terms of service may entail for the researcher. Second, we present a scraping routine for public Facebook posts, and discuss some technical adjustments that can be performed for the data to be ethically and legally acceptable. The code employs screen scraping to collect the list of reactions to a Facebook public post, and performs a one-way cryptographic hash function on the users' identifiers to pseudonymize their personal information, while still keeping them traceable within the data. This article contributes to the debate around freedom of internet research and the ethical concerns that might arise by scraping data from the social web.