TextRWeb: Large-Scale Text Analytics with R on the Web

TextRWeb: Large-Scale Text Analytics with R on the Web
复制标题

TextRWeb:在网络上使用 R 进行大规模文本分析

DOI:
--
复制
发表时间:
2014
期刊:
Extreme Science and Engineering Discovery Environment
影响因子:
--
通讯作者:
Beth Plale
Beth Plale
中科院分区:
--
文献类型:
--
作者:
Guangchen Ruan;Hui Zhang;E. Wernert;Beth Plale

文献摘要

被引文献

相似文献

随着数字数据源的数量和规模的增长,它们为通过文本挖掘,NLP和其他文本分析技术进行计算调查提供了机会。R是一种流行且功能强大的文本分析工具;然而,它需要并行运行,并需要特殊处理以保护受版权保护的内容免受完全访问(消费)。HathiTrust研究中心(HTRC)目前拥有1100万册(书籍),其中700万册有版权。在本文中,我们提出了HTRC TextRWeb,一个交互式的R软件环境,采用复杂性隐藏接口和自动代码生成,允许大规模的文本分析在一个非消费性的手段。对于HathiTrust数字图书馆中受版权保护的数据的主要测试案例,TextRWeb允许我们编码,编辑和提交由一系列交互式Web用户界面授权的文本分析方法。所有这些方法联合收割机揭示了一个新的互动模式,大规模的文本分析在网络上。
As digital data sources grow in number and size, they pose an opportunity for computational investigation by means of text mining, NLP, and other text analysis techniques. R is a popular and powerful text analytics tool; however, it needs to run in parallel and requires special handling to protect copyrighted content against full access (consumption). The HathiTrust Research Center (HTRC) currently has 11 million volumes (books) where 7 million volumes are copyrighted. In this paper we propose HTRC TextRWeb, an interactive R software environment which employs complexity hiding interfaces and automatic code generation to allow large-scale text analytics in a non-consumptive means. For our principal test case of copyrighted data in HathiTrust Digital Library, TextRWeb permits us to code, edit, and submit text analytics methods empowered by a family of interactive web user interfaces. All these methods combine to reveal a new interactive paradigm for large-scale text analytics on the web.