Defining logical domains in a web site

Defining logical domains in a web site
复制标题

在网站中定义逻辑域

DOI:
10.1145/336296.336345
复制
发表时间:
2000
期刊:
--
影响因子:
--
通讯作者:
H. Takano
H. Takano
中科院分区:
--
文献类型:
--
作者:
Wen;O. Kolak;Q. Vu;H. Takano

文献摘要

被引文献

相似文献

每个URL标识一个唯一的Web页面;因此,它被视为用于组织Web查询结果的自然选择。Web搜索结果可以按域分组并作为集群呈现给用户以便于可视化。然而,它有一个缺点:处理大型网站,如Geocities,W3C,ndwww.cs.umd.edu。大型Web站点往往会产生许多匹配项,从而产生一些大型、扁平结构化和无组织的集群。事实上,这些站点包含其他实体(如项目和人员)的Web站点。这些网站中的许多页面实际上本身就是“逻辑域”。例如,大学项目的网站或W3C的XML部分可以被视为“逻辑域”。本文提出了逻辑域的概念,它是相对于简单地用域名来标识的物理域而言的。我们已经开发和实现了一套规则的基础上,链接结构,路径信息,文档元数据,和引文,以确定逻辑域入口页面和它们相应的边界。在真实的Web数据上的实验验证了该方法的有效性。
Each URL identifies a unique Web page; thus, it is viewed as a natural choice to use for organizing Web query results. Web search results may be grouped by domain and presented to users as clusters for ease of visualization. However, it has a drawback: dealing with large Web sites, such as Geocities, W3C ,a ndwww.cs.umd.edu. Large Web sites tend to yield many matches that leads to a few large, flat structured, and unorganized clusters. As a matter of fact, these sites contain Web sites of other entities, such as projects and people. Many pages in these sites are actually “logical domains” by themselves. For example, Web sites for projects at a university or the XML section at W3C could be viewed as “logical domains”. In this paper, we propose the concept of logical domain with respect to physical domainwhich is identified simply by domain name. We have developed and implemented a set of rules based on link structure, path information, document metadata, and citation to identify logical domain entry pages and their corresponding boundaries. Experiments on real Web data have been conducted to validate the usefulness of this technique.