Automatic Data Extraction from Lists and Tables in Web Sources
Automatic Data Extraction from Lists and Tables in Web Sources
复制标题
从 Web 源中的列表和表格自动提取数据
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
M. Rey
中科院分区:
文献类型:
--
作者:
Kristina Lerman;Craig A. Knoblock;Steven N. Minton;M. Rey
We describe a technique for extracting data from lists and tables and grouping it by rows and columns. This is done completely automatically, using only some very general assumptions about the structure of the list. We have developed a suite of unsupervised learning algorithms that induce the structure of lists by exploiting the regularities both in the format of the pages and the data contained in them. Among the tools used are AutoClass for automatic classification of data and grammar induction of regular languages. The approach was tested on 14 Web sources providing diverse data types, and we found that for 10 of these sources we were able to correctly find lists and partition the data into columns and rows.