Design and implementation of a scalable distributed web crawler based on Hadoop
Design and implementation of a scalable distributed web crawler based on Hadoop
复制标题
基于Hadoop的可扩展分布式网络爬虫的设计与实现
DOI:
10.1109/icbda.2017.8078691
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
T. Zhang
中科院分区:
文献类型:
--
作者:
Yuliang Shi;T. Zhang
In this article, an efficient and scalable distributed web crawler system based on Hadoop will be design and implement. In the paper, firstly the application of cloud computing in reptile field is introduced briefly, and then according to the current status of the crawler system, the specific use of Hadoop distributed and cloud computing features detailed design of a highly scalable crawler system, and finally the system Data statistics, under the same conditions, compared with the existing mature system, it is clear that the superiority of distributed web crawler. This advantage in the context of large data era of massive data is particularly important to climb.