A free database of university web links: data collection issues

A free database of university web links: data collection issues
复制标题

大学网络链接的免费数据库:数据收集问题

DOI:
--
复制
发表时间:
2002
期刊:
影响因子:
--
通讯作者:
M. Thelwall
M. Thelwall
中科院分区:
--
文献类型:
--
作者:
M. Thelwall

文献摘要

被引文献

相似文献

本文描述了一组免费的数据库,其中包含来自选定国家的大学网站的链接结构,由专业的信息科学网络爬虫创建。随着信息和计算机科学家对网络链接的兴趣日益浓厚,这是一种为研究提供原始数据的尝试,这些数据不依赖于商业搜索引擎的不透明技术。还提供了基本的查询工具。还讨论了有关运行准确的网络爬虫的关键问题。还可以访问通常隐藏的爬行程序停止列表,目的是使爬行过程更加透明。讨论了建立这样一个列表的必要性,得出的结论是,由于数据库生成的网络区域的存在以及镜像现象的扩散,全自动爬行在社会上或经验上都不是可取的。
This paper describes a free set of databases of the link structures of the university web sites from a selection of countries, as created by a specialist information science web crawler. With the increasing interest in web links by information and computer scientists this is an attempt to make available raw data for research that is not reliant upon the opaque techniques of commercial search engines. Basic tools for querying are also provided. The key issues concerning running an accurate web crawler are also discussed. Access is also given to the normally hidden crawler stop list with the aim of making the crawl process more transparent. The necessity of having such a list is discussed, with the conclusion that fully automatic crawling is not socially or empirically desirable because of the existence of database-generated areas of the web and the proliferation of the phenomenon of mirroring.