Extracting the Latent Hierarchical Structure of Web Documents

Extracting the Latent Hierarchical Structure of Web Documents
复制标题

提取 Web 文档的潜在层次结构

DOI:
10.1007/978-3-642-01350-8_28
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
A. Rafea
A. Rafea
中科院分区:
--
文献类型:
--
作者:
Michael A. El;S. El;A. Rafea

文献摘要

被引文献

相似文献

文档的层次结构在理解其内容之间的关系方面起着重要作用。然而,这样的结构并不总是通过可用的html分层标记在web文档中显式地表示。然而,标题通常在呈现方面与文档中的“正常”文本区分开,从而提供了人类读者可辨别的隐含结构。因此,对于需要在层次级别上操作的应用程序来说,一个重要的预处理步骤是提取隐式表示的层次结构。本文提出了一种利用多种视觉信息进行航向检测和航向水平检测的算法。评估该算法的结果也报告。
The hierarchical structure of a document plays an important role in understanding the relationships between its contents. However, such a structure is not always explicitly represented in web documents through available html hierarchical tags. Headings however, are usually differentiated from ‘normal’ text in a document in terms of presentation thus providing an implicit structure discernable by a human reader. As such, an important pre-processing step for applications that need to operate on the hierarchical level is to extract the implicitly represented hierarchal structure. In this paper, an algorithm for heading detection and heading level detection which makes use of various visual presentations is presented. Results of evaluating this algorithm are also reported.