Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining

Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining
复制标题

使用 R 自动数据收集:网页抓取和文本挖掘实用指南

DOI:
10.1002/9781118834732
复制
发表时间:
2014
期刊:
Information, Communication & Society
影响因子:
--
通讯作者:
Dominic Nyhuis
Dominic Nyhuis
中科院分区:
--
文献类型:
--
作者:
Simon Munzert;C. Rubba;Peter Meiner;Dominic Nyhuis

文献摘要

被引文献

相似文献

面向 R 初学者和经验丰富的用户的网络抓取和文本挖掘实践指南介绍了网络和数据库主要架构的基本概念,涵盖 HTTP、HTML、XML、JSON、SQL。提供查询 Web 文档和数据集(XPath 和正则表达式)的基本技术。提供了一系列广泛的练习来指导读者完成每种技术。探索监督和非监督技术以及数据抓取和文本管理等高级技术。贯穿全文的案例研究以及所介绍的每种技术的示例。本书中的 R 代码和练习解决方案在支持网站上提供。
A hands on guide to web scraping and text mining for both beginners and experienced users of RIntroduces fundamental concepts of the main architecture of the web and databases and covers HTTP, HTML, XML, JSON, SQL. Provides basic techniques to query web documents and data sets (XPath and regular expressions). An extensive set of exercises are presentedto guide the reader through each technique. Explores both supervised and unsupervised techniques as well as advanced techniques such as data scraping and text management. Case studies are featured throughout along with examples for each technique presented. R code and solutionsto exercises featuredin the book are provided ona supporting website.