Web Data Extraction Based on XBRL-GL Taxonomy

Web Data Extraction Based on XBRL-GL Taxonomy
复制标题

基于XBRL-GL分类法的Web数据提取

DOI:
10.1109/apcip.2009.97
复制
发表时间:
2009
期刊:
2009 Asia-Pacific Conference on Information Processing
影响因子:
--
通讯作者:
Jin
Jin
中科院分区:
--
文献类型:
--
作者:
Hanyang Luo;Jin

文献摘要

被引文献

相似文献

Web已经成为连接各种信息资源的最重要的连接之一。最令人感兴趣的挑战是如何从大量的Web页面中提取重要数据,并将其转换为更具结构化、标准化和语义性的信息,以便利用数据库、数据仓库等领域的成熟技术进行查询和分析。本文基于XBRL-GL分类法,将数据抽取技术与XBRL技术相结合,设计了一个包装器生成器。该包装器可以根据对HTML文档结构的分析将其转换为XML格式,然后使用XPath来定位数据。这样,我们就可以准确地提取数据,并以标准的形式存储。
The Web has become one of the most important connections to various information resources. The most interesting challenge is how to extract important data from a large number of web pages and transform them to more structural, standard and semantic information, which can be queried and analyzed by using matured techniques in database, data warehouse and other fields. We design a wrapper generator by combining the data extraction technique with XBRL technology based on XBRL-GL taxonomy. The wrapper can transform HTML documents to XML forms according to the analysis of HTML document structure, and then use XPath to locate the data. In this way, we can extract the data accurately and store them in a standard form.