Finding and Extracting Data Records from Web Pages

Finding and Extracting Data Records from Web Pages
复制标题

DOI:
10.1007/s11265-008-0270-y
复制
发表时间:
2010-04-01
影响因子:
1.8
通讯作者:
Cacheda, Fidel
Cacheda, Fidel
中科院分区:
计算机科学4区
文献类型:
--
作者:
Alvarez, Manuel;Pan, Alberto;Cacheda, Fidel

文献摘要

被引文献

相似文献

许多HTML页是由软件程序通过查询一些底层数据库,然后用数据填充模板来生成的。在这些情况下,关于数据结构的元信息会丢失,因此自动化软件程序无法以强大的方式处理这些数据,例如来自数据库的信息。我们提出了一套新颖的技术来检测网页中的结构化记录并提取构成它们的数据值。我们的方法只需要一个输入页面。它从识别页面中感兴趣的数据区域开始。然后,通过使用将页面的DOM树中的相似子树分组的聚类方法将其划分为记录。最后,采用基于多串对齐的方法提取数据记录的属性。我们用大量真实的网络资源测试了我们的技术,获得了很高的查准率和查全率。
Many HTML pages are generated by software programs by querying some underlying databases and then filling in a template with the data. In these situations the metainformation about the data structure is lost, so automated software programs cannot process these data in such powerful manners as information from databases. We propose a set of novel techniques for detecting structured records in a web page and extracting the data values that constitute them. Our method needs only an input page. It starts by identifying the data region of interest in the page. Then it is partitioned into records by using a clustering method that groups similar subtrees in the DOM tree of the page. Finally, the attributes of the data records are extracted by using a method based on multiple string alignment. We have tested our techniques with a high number of real web sources, obtaining high precision and recall values.