Web Data Extraction Based on XBRL-GL Taxonomy
Web Data Extraction Based on XBRL-GL Taxonomy
复制标题
基于XBRL-GL分类法的Web数据提取
DOI:
10.1109/apcip.2009.97
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
Jin
中科院分区:
文献类型:
--
作者:
Hanyang Luo;Jin
The Web has become one of the most important connections to various information resources. The most interesting challenge is how to extract important data from a large number of web pages and transform them to more structural, standard and semantic information, which can be queried and analyzed by using matured techniques in database, data warehouse and other fields. We design a wrapper generator by combining the data extraction technique with XBRL technology based on XBRL-GL taxonomy. The wrapper can transform HTML documents to XML forms according to the analysis of HTML document structure, and then use XPath to locate the data. In this way, we can extract the data accurately and store them in a standard form.