Browser Record and Replay as a Building Block for End-User Web Automation Tools

Browser Record and Replay as a Building Block for End-User Web Automation Tools
复制标题

浏览器记录和重放作为最终用户 Web 自动化工具的构建块

DOI:
10.1145/2740908.2742849
复制
发表时间:
2015
期刊:
Proceedings of the 24th International Conference on World Wide Web
影响因子:
--
通讯作者:
Sumit Gulwani
Sumit Gulwani
中科院分区:
--
文献类型:
--
作者:
Sarah E. Chasins;Shaon Barman;Rastislav Bodík;Sumit Gulwani

文献摘要

被引文献

相似文献

要为最终用户的演示(PBD)Web刮擦工具构建编程,一个需要两个中心组件:列表查找器以及一个记录和重播工具。列表查找器从网页中提取逻辑表。记录和重播(R+R)系统记录用户与网页的交互,并以编程方式进行重新播放。研究界已经在清单查找中投入了大量工作 - 各种称为包装器归纳,结构化数据提取和模板检测。相比之下,研究人员在很大程度上考虑了浏览器R+R问题,直到最近,网页复杂性和交互性开始上升。 We argue that the increase in interactivity necessitates the use of new, more robust R+R approaches, which will facilitate the PBD web tools of the future.由于稳健的R+R很难构建和理解,因此我们认为工具开发人员需要一个可以将其视为黑匣子的R+R层。我们设计了一个易于使用的API,允许程序员使用甚至自定义R+R,而无需了解R+R内部词。我们已在Ringer(我们的强大的R+R工具)中实例化API。我们使用API​​实现WebCombine,这是一种PBD刮擦工具。 WebCombine用户演示了如何收集关系数据集的第一行,该工具收集所有剩余行。 WebCombine使用Ringer API来处理页面之间的导航,从而使用户可以从现代互动较重的页面中刮擦。 We demonstrate WebCombine by collecting a 3,787,146 row dataset from Google Scholar that allows us to explore the relationship between researchers' years of experience and their papers' citation counts.
To build a programming by demonstration (PBD) web scraping tool for end users, one needs two central components: a list finder, and a record and replay tool. A list finder extracts logical tables from a webpage. A record and replay (R+R) system records a user's interactions with a webpage, and replays them programmatically. The research community has invested substantial work in list finding --- variously called wrapper induction, structured data extraction, and template detection. In contrast, researchers largely considered the browser R+R problem solved until recently, when webpage complexity and interactivity began to rise. We argue that the increase in interactivity necessitates the use of new, more robust R+R approaches, which will facilitate the PBD web tools of the future. Because robust R+R is difficult to build and understand, we argue that tool developers need an R+R layer that they can treat as a black box. We have designed an easy-to-use API that allows programmers to use and even customize R+R, without having to understand R+R internals. We have instantiated our API in Ringer, our robust R+R tool. We use the API to implement WebCombine, a PBD scraping tool. A WebCombine user demonstrates how to collect the first row of a relational dataset, and the tool collects all remaining rows. WebCombine uses the Ringer API to handle navigation between pages, enabling users to scrape from modern, interaction-heavy pages. We demonstrate WebCombine by collecting a 3,787,146 row dataset from Google Scholar that allows us to explore the relationship between researchers' years of experience and their papers' citation counts.