A method for language-specific Web crawling and its evaluation

A method for language-specific Web crawling and its evaluation
复制标题

一种特定语言的Web爬行方法及其评估

DOI:
10.1002/scj.20693
复制
发表时间:
2007
期刊:
Syst. Comput. Jpn.
影响因子:
--
通讯作者:
M. Kitsuregawa
M. Kitsuregawa
中科院分区:
--
文献类型:
--
作者:
T. Tamura;Kulwadee Somboonviwat;M. Kitsuregawa

文献摘要

被引文献

相似文献

许多国家已经建立了网络存档项目,旨在长期保存网络信息,这些信息现在被认为是文化和社会方面的珍贵信息。然而,由于其无国界的特点,网络对全面收集源自特定国家或文化的信息构成了障碍。本文提出了一种有效的方法,有选择地收集网页写在一个特定的语言。首先,从一个大的爬行获得的真实的Web数据的语言图分析进行,以获得爬行的指导方针,它利用每个Web服务器的语言属性。然后,该指南形成为链接选择策略的几个变体。基于模拟的评估表明,策略之一,仔细接受新发现的Web服务器,显示上级的结果,收获率/覆盖率和运行时效率。© 2007 Wiley Periodicals,Inc. Syst Comp Jpn,38(2):10-20,2007;在线发表于Wiley InterScience(www.interscience.wiley.com)。DOI 10.1002/scj.20693
Many countries have created Web archiving projects aiming at long-term preservation of Web information, which is now considered precious in cultural and social aspects. However, because of its borderless character, the Web poses obstacles to comprehensively gathering information originating in a specific nation or culture. This paper proposes an efficient method for selectively collecting Web pages written in a specific language. First, a linguistic graph analysis of real Web data obtained from a large crawl is conducted in order to derive a crawling guideline, which makes use of language attributes per Web server. The guideline then is formed into a few variations of link selection strategies. Simulation-based evaluation reveals that one of the strategies, which carefully accepts newly discovered Web servers, shows superior results in terms of harvest rate/coverage and runtime efficiency. © 2007 Wiley Periodicals, Inc. Syst Comp Jpn, 38(2): 10–20, 2007; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/scj.20693