tidytext: Text Mining and Analysis Using Tidy Data Principles in R

tidytext: Text Mining and Analysis Using Tidy Data Principles in R
复制标题

tidytext:在 R 中使用 Tidy 数据原理进行文本挖掘和分析

DOI:
--
复制
发表时间:
2016
影响因子:
--
通讯作者:
David Robinson
David Robinson
中科院分区:
--
文献类型:
--
作者:
Julia Silge;David Robinson

文献摘要

被引文献

相似文献

整洁的数据集允许使用一套标准的“整洁”工具进行操作,包括流行的软件包,如dplyr (Wickham, Francois和RStudio 2015), ggplot2 (Wickham, Chang和RStudio 2016)和broom (Robinson et al. 2015)。然而,这些工具还没有能够流畅地处理文本数据和自然语言处理工具的基础设施。在开发这个包的过程中,我们提供了功能和支持数据集,以允许文本与整洁格式之间的转换,并在整洁工具和现有文本挖掘包之间无缝切换。
Tidy data sets allow manipulation with a standard set of “tidy” tools, including popular packages such as dplyr (Wickham, Francois, and RStudio 2015), ggplot2 (Wickham, Chang, and RStudio 2016), and broom (Robinson et al. 2015). These tools do not yet, however, have the infrastructure to work fluently with text data and natural language processing tools. In developing this package, we provide functions and supporting data sets to allow conversion of text to and from tidy formats, and to switch seamlessly between tidy tools and existing text mining packages.