Automatic Data Extraction from Lists and Tables in Web Sources

Automatic Data Extraction from Lists and Tables in Web Sources
复制标题

从 Web 源中的列表和表格自动提取数据

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
M. Rey
M. Rey
中科院分区:
--
文献类型:
--
作者:
Kristina Lerman;Craig A. Knoblock;Steven N. Minton;M. Rey

文献摘要

被引文献

相似文献

我们描述了一种从列表和表格中提取数据并按行和列进行分组的技术。这是完全自动完成的,只使用了一些关于列表结构的非常一般的假设。我们已经开发了一套无监督学习算法,通过利用页面格式和其中包含的数据来诱导列表的结构。其中使用的工具是自动分类的数据和规则语言的语法归纳AutoClass。该方法在14个提供不同数据类型的Web源上进行了测试,我们发现对于其中的10个源,我们能够正确地找到列表并将数据划分为列和行。
We describe a technique for extracting data from lists and tables and grouping it by rows and columns. This is done completely automatically, using only some very general assumptions about the structure of the list. We have developed a suite of unsupervised learning algorithms that induce the structure of lists by exploiting the regularities both in the format of the pages and the data contained in them. Among the tools used are AutoClass for automatic classification of data and grammar induction of regular languages. The approach was tested on 14 Web sources providing diverse data types, and we found that for 10 of these sources we were able to correctly find lists and partition the data into columns and rows.