Towards data extraction of dynamic content from JavaScript Web applications

Towards data extraction of dynamic content from JavaScript Web applications
复制标题

从 JavaScript Web 应用程序中提取动态内容的数据

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Information Networking
影响因子:
--
通讯作者:
Korawit Prutsachainimmit
Korawit Prutsachainimmit
中科院分区:
--
文献类型:
--
作者:
W. Nadee;Korawit Prutsachainimmit

文献摘要

被引文献

相似文献

万维网和社交媒体中的大量数据为企业和组织提供了获得有效运营的重要价值的机会。因此,Web数据抽取已经成为收集和翻译半结构化文档为有价值的信息的重要工具。然而,主要挑战之一是处理Web文档的变化,特别是JavaScript Web开发技术的出现,它极大地影响了嵌入和呈现网页数据的方式。在本文中,我们提出了一个新的Web数据抽取系统,旨在提取数据从JavaScript Web应用程序的设计和实现。该系统通过定义数据抽取规则和数据转换模式,使用户能够从在线Web文档中选择有价值的数据。抽取引擎自动地将半结构化数据抓取并转换为关系数据。初步评估结果表明,我们提出的系统已经成功地从现代JavaScript Web应用程序中提取数据。
An enormous data in World Wide Web and social media has open opportunities for business and organization to get the significant value that leads to efficient operations. As a result, Web Data Extraction has become an important tool for gathering and translating semi-structured documents into valuable information. However, one of the major challenges is dealing with changes from Web documents, especially emerging of JavaScript Web development technology that has significantly affected the way to embed and rendering data of Web pages. In this paper, we propose a design and implementation of a new Web Data Extraction system that aims for extracting data from JavaScript Web applications. The proposed system enables users to select valuable data from online Web documents by defining data extraction rules and data transformation patterns. The extraction engine automatically scrapes and transforms semi-structure data into relational data. The preliminary evaluation results showed that our proposed system has successfully extract data from modern JavaScript Web applications.