Design and implementation of a scalable distributed web crawler based on Hadoop

Design and implementation of a scalable distributed web crawler based on Hadoop
复制标题

基于Hadoop的可扩展分布式网络爬虫的设计与实现

DOI:
10.1109/icbda.2017.8078691
复制
发表时间:
2017
期刊:
2017 IEEE 2nd International Conference on Big Data Analysis (ICBDA)(
影响因子:
--
通讯作者:
T. Zhang
T. Zhang
中科院分区:
--
文献类型:
--
作者:
Yuliang Shi;T. Zhang

文献摘要

被引文献

相似文献

本文将设计并实现一个基于Hadoop的高效、可扩展的分布式网络爬虫系统。论文首先简单介绍了云计算在爬虫领域的应用,然后根据爬虫系统的现状,具体利用Hadoop分布式和云计算的特点详细设计了一个高扩展性的爬虫系统,最后对系统的数据进行统计,在同等条件下,与现有的成熟系统进行比较,很明显分布式网络爬虫的优越性。这一优势在大数据时代海量数据的攀爬背景下显得尤为重要。
In this article, an efficient and scalable distributed web crawler system based on Hadoop will be design and implement. In the paper, firstly the application of cloud computing in reptile field is introduced briefly, and then according to the current status of the crawler system, the specific use of Hadoop distributed and cloud computing features detailed design of a highly scalable crawler system, and finally the system Data statistics, under the same conditions, compared with the existing mature system, it is clear that the superiority of distributed web crawler. This advantage in the context of large data era of massive data is particularly important to climb.