Extracting the Latent Hierarchical Structure of Web Documents
Extracting the Latent Hierarchical Structure of Web Documents
复制标题
提取 Web 文档的潜在层次结构
DOI:
10.1007/978-3-642-01350-8_28
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
A. Rafea
中科院分区:
文献类型:
--
作者:
Michael A. El;S. El;A. Rafea
The hierarchical structure of a document plays an important role in understanding the relationships between its contents. However, such a structure is not always explicitly represented in web documents through available html hierarchical tags. Headings however, are usually differentiated from ‘normal’ text in a document in terms of presentation thus providing an implicit structure discernable by a human reader. As such, an important pre-processing step for applications that need to operate on the hierarchical level is to extract the implicitly represented hierarchal structure. In this paper, an algorithm for heading detection and heading level detection which makes use of various visual presentations is presented. Results of evaluating this algorithm are also reported.