Using common hypertext links to identify the best phrasal description of target web documents

Using common hypertext links to identify the best phrasal description of target web documents
复制标题

使用常见的超文本链接来识别目标 Web 文档的最佳短语描述

DOI:
--
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
E. Amitay
E. Amitay
中科院分区:
--
文献类型:
--
作者:
E. Amitay

文献摘要

被引文献

相似文献

本文介绍了前人的工作,研究并比较了Web文档中的单词分布与“正常”平面文本中的单词分布。根据这项研究的发现,传统的信息检索技术不能像用于“普通”文本集合(例如新闻文章)那样用于网络搜索目的。然后,基于这些相同的发现,我将描述一种新的文档描述模型,该模型利用了传统技术忽略的Web上提供的有价值的锚文本信息。The Problem Amitay(1997)通过对1000个网页的语料库分析发现,在专门为网络编写的文档(主页)中的词汇分布与在“正常”英语语料库(英国国家语料库100,000,00字)中观察到的词汇分布显著不同。例如,在Web Documents集合中,有一些不包含动词或限定词(即“the”、“a”等)的HTML文件。尽管它们有60多个单词(不包括HTML标签)。虽然“the”一词在整个网络集合中约占3%,但在英语集合(BNC)中,它约占7%。这项研究还发现,在Web上,人们编写文档时有一种惯例:在锚文本(即<a href>文本</a>)中有一定数量的词用于描述其他目标文档,并且在短语、句子或列表中“突出显示”这些词是有语言惯例的。以前的工作和解决方案本节介绍以前对在网上查找信息问题提出的解决方案,其中考虑到信息的结构和网络作者提供的其他元数据。这些解决方案按时间顺序提出,有趣的是,随着时间的推移,解决方案越来越依赖于网页作者提供的信息(例如链接结构、锚文本等)。由于本文建议使用锚和链接结构中嵌入的信息,因此选择下面描述的研究以显示过去使用此类信息的趋势。Frei和Stieger(1992)描述了一种使用超文本链接的语义内容进行检索的方法。他们提出了一种索引算法,该算法利用文档的文本和链接内容。链接的内容被标记为“参考的”或“语义的”。还为文本内容和指向/节点关系标记语义链接。McBryan(1994)建议可以对标题、参考超文本或URL名称字符串的组成部分进行搜索。在他的系统中,他用其锚和Page Plus标题为每个URL编制索引
This paper describes previous work which studied and compared the distribution of words in web documents with the distribution of words in "normal" flat texts. Based on the findings from this study it is suggested that the traditional IR techniques cannot be used for web search purposes the same way they are used for "normal" text collections, e.g. news articles. Then, based on these same findings, I will describe a new document description model which exploits valuable anchor text information provided on the web that is ignored by the traditional techniques. The problem Amitay (1997) has found, through a corpus analysis of a 1000 web pages that the lexical distribution in documents which were written especially for the web (home pages), is significantly different than the lexical distribution observed in a corpus of "normal" English language (the British National Corpus 100,000,00 words). For example, in the web documents collection there were some HTML files which contained no verb or determiner (i.e. "the", "a", etc.) although there are more than 60 words in them (excluding the HTML tags). While the word "the" comprised around 3% of the whole web collection, in the English language collection (BNC) it comprises about 7%. This study also found that on the web there is a convention with which people write their documents: there is a certain number of words used to describe other target documents in the anchor text (i.e. <a href="">text</a>), and there is a linguistic convention in "highlighting" these words in the phrase, sentence or list. Previous work and solutions This section describes previous suggested solutions to the problem of finding information on the web, taking into account its structure and the additional meta-data provided by web authors. The solutions are presented in a chronological order and it is interesting to note that, through time, solutions rely more and more on the information provided by the authors of the web pages (e.g. link structure, anchor text, etc.). Since this paper suggests using the information embedded in the anchors and the link structure, the studies described below were chosen in order to show past trends in using such information. Frei and Stieger (1992) describe a way for using the semantic content of hypertext links for retrieval. They present an indexing algorithm which makes use of the document's text and link content. The content of the link is marked as being "referential" or "semantic". Semantic links are further marked for textual content and pointing/node relations. McBryan (1994) suggests that searches can be performed on titles, reference hypertext, or within components of URL name strings. In his system he indexes each URL with its anchor and title of page plus