EXPERIMENTING WITH RULE LEARNING FOR INFORMATION EXTRACTION FROM HTML

EXPERIMENTING WITH RULE LEARNING FOR INFORMATION EXTRACTION FROM HTML
复制标题

试验从 HTML 中提取信息的规则学习

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
A. Bădică
A. Bădică
中科院分区:
--
文献类型:
--
作者:
C. Bǎdicǎ;A. Bădică

文献摘要

被引文献

相似文献

Web是一个不断增长的信息库,具有跨越许多应用领域的丰富语义结构。然而,网络主要是为人类消费而设计的,而不是自动处理。这是实现信息搜索、过滤和提取等任务自动化的主要障碍。在此背景下,本文的目的是提出一种从表示产品信息表的HTML信息源中学习规则来提取产品信息的技术。该技术利用了这样一个事实,即表示某个生产商的产品信息的网页是从生产商数据库中动态生成的,因此它们呈现出统一的结构。因此,虽然提取任务是由人类用户针对少数信息项手动执行的,但通用归纳学习者可以学习将被进一步应用于当前和其他产品信息表以自动提取其他项的提取规则。学习算法的输入是定义了HTML树节点类型和它们之间的关系的HTML文档树的关系描述。通过适当的实例、实验结果和软件工具对该方法进行了验证。
The Web is a continuously growing information repository with a rich semantic structure that spans many application areas. The Web, however, has been designed primarily for human consumption rather than automated processing. This is a major obstacle for automating tasks like information searching, filtering and extraction. In this context, the aim of the paper is to present a technique for learning rules to extract product information from HTML information sources that represent product information sheets. The technique exploits the fact that the Web pages that represent product information of a certain producer are generated on the fly from the producer database and therefore they exhibit uniform structures. Consequently, while the extraction task is executed manually for a few information items by a human user, a general-purpose inductive learner can learn extraction rules that will be further applied to the current and other product information sheets to automatically extract other items. The input to the learning algorithm is a relational description of the HTML document tree that defines the HTML tree nodes types and the relationships between them. The approach is demonstrated with appropriate examples, experimental results, and software tools.