Applying site information to information retrieval from the Web

Applying site information to information retrieval from the Web
复制标题

将站点信息应用于 Web 信息检索

DOI:
10.1109/wise.2002.1181646
复制
发表时间:
2002
期刊:
Proceedings of the Third International Conference on Web Information Systems Engineering, 2002. WISE 2002.
影响因子:
--
通讯作者:
M. Kitsuregawa
M. Kitsuregawa
中科院分区:
--
文献类型:
--
作者:
Yasuhito Asano;H. Imai;Masashi Toyoda;M. Kitsuregawa

文献摘要

被引文献

相似文献

近年来,已经开发了几种使用关于Web链接的信息的信息检索方法,例如HITS和拖网。为了分析划分为每个Web站点内部的链接(本地链接)和Web站点之间的链接(全局链接)以用于信息检索的Web链接,需要适当的Web站点模型。在现有的研究中,Web服务器被用作Web站点的模型。当一个Web站点对应于一个服务器时,这种想法工作得相对较好,就像公共Web站点的情况一样,但是当多个Web站点对应于一个服务器时,这种想法工作得很差,就像租用Web服务器上的私有Web站点的情况一样。我们提出了一种新的模型的网站,“基于目录的网站”,以处理典型的私人网站,和一种方法来识别他们使用的URL和Web链接的信息。我们验证了该方法可以以超过110,000台服务器中66%的速度大约识别每个服务器是否有多个基于目录的网站,并通过使用jp的计算实验提取超过500,000个基于目录的网站和400万个全局链接域名URL和Web链接数据包含超过2300万个URL和1亿个Web链接,2000年7月至8月,由丰田和Kitsuregawa收集。我们还提出了一个新的框架,基于Web链接的信息检索,使用基于目录的网站和全球链接,而不是网页和整个Web链接分别,并检查我们的框架的有效性进行比较,拖网在我们的框架上的一个现有的框架。
In recent years, several information retrieval methods using information about Web-links have been developed, such as HITS and trawling. In order to analyze Web-links dividing into links inside each Web site (local-links) and links between Web sites (global-links)for information retrieval, a proper model of the Web site is required. In existing research, a Web server is used as a model of the Web site. This idea works relatively well when a Web site corresponds to a server, as is the case for public Web sites, but works poorly when multiple Web sites correspond to a server, as is the case for private Web sites on rental Web servers. We propose a new model of the Web site, "directory-based site", to handle typical private sites, and a method to identify them using information about the URL and Web-links. We verify the method can approximately identify, at a rate of 66% of over 110,000 servers, whether each server has multiple directory-based sites or not, and extract over 500,000 directory-based sites and 4 million global-links by computational experiments using jp-domain URLs and Web-link data contains over 23 million URLs and 100 million Web-links, collected from July to August 2000, by Toyoda and Kitsuregawa. We also propose a new framework of Web-link based information retrieval that uses directory-based sites and global-links instead of Web pages and whole Web-links respectively, and examine the effectiveness of our framework by comparing a result of trawling on our framework to one on the existing framework.