Using the Wayback Machine to Mine Websites in the Social Sciences: A Methodological Resource

Using the Wayback Machine to Mine Websites in the Social Sciences: A Methodological Resource
复制标题

DOI:
10.1002/asi.23503
复制
发表时间:
2016-08-01
影响因子:
3.5
通讯作者:
Shapira, Philip
Shapira, Philip
中科院分区:
管理学3区
文献类型:
--
作者:
Arora, Sanjay K.;Li, Yin;Shapira, Philip

文献摘要

被引文献

相似文献

网站为开发和分析各种社会科学现象的信息提供了一个不引人注目的数据源。在本文中,我们为社会科学家提供了一种方法资源,他们希望使用非结构化的基于网络的文本来扩展他们的工具包,特别是使用Wayback Machine来访问历史网站数据。在提供了使用时光机的现有研究的文献综述之后,我们提出了分析人员如何使用存档网站设计研究项目的逐步描述。我们以一个项目为例,该项目分析了300家美国绿色产品行业中小企业的创新活动和战略指标。我们提出了访问历史Wayback网站数据的六个步骤:(a)采样,(b)组织和定义网络抓取的边界,(c)抓取,(d)网站变量操作,(e)与其他数据源集成,以及(f)分析。虽然我们的例子借鉴了绿色产品行业中特定类型的公司,但该方法可以推广到其他研究领域。在讨论使用Wayback Machine的限制和好处时,我们注意到机器和人的努力对于从存档的web信息中开发高质量的数据集是必不可少的。
Websites offer an unobtrusive data source for developing and analyzing information about various types of social science phenomena. In this paper, we provide a methodological resource for social scientists looking to expand their toolkit using unstructured web-based text, and in particular, with the Wayback Machine, to access historical website data. After providing a literature review of existing research that uses the Wayback Machine, we put forward a step-by-step description of how the analyst can design a research project using archived websites. We draw on the example of a project that analyzes indicators of innovation activities and strategies in 300 U.S. small- and medium-sized enterprises in green goods industries. We present six steps to access historical Wayback website data: (a) sampling, (b) organizing and defining the boundaries of the web crawl, (c) crawling, (d) website variable operationalization, (e) integration with other data sources, and (f) analysis. Although our examples draw on specific types of firms in green goods industries, the method can be generalized to other areas of research. In discussing the limitations and benefits of using the Wayback Machine, we note that both machine and human effort are essential to developing a high-quality data set from archived web information.