Text Analysis with R for Students of Literature

Text Analysis with R for Students of Literature
复制标题

文学专业学生使用 R 进行文本分析

DOI:
--
复制
发表时间:
2016
影响因子:
1.4
通讯作者:
L. Lei
L. Lei
中科院分区:
人文科学4区
文献类型:
--
作者:
L. Lei

文献摘要

被引文献

相似文献

文学领域的研究在很大程度上是基于定性的,该领域的人们有时可能会质疑在文学研究中使用定量和计算方法的意义。作者在本书的序言中讨论了这个问题,认为计算方法的结果可以为“推测方法”的结果提供补充证据,并可能得出新的发现(第viii页)。本书是人文社会科学系列定量方法之一。它的目的是提供文学,或更普遍的人文主义,研究人员的定量和计算方法研究的文本(第七)。沿着这些方法,这本书讨论了R的基础知识,R是一种免费但功能强大的语言和文本分析和统计工具。由于大多数目标读者可能没有编程经验,本书假设读者没有编程背景知识。话虽如此,这本书在结构上不同于其他类似目的的优秀书籍,如Baayen(2008),Gries(2009)和Gries(2013)。它不是从简单介绍编程语言的基本数据类型开始,然后一个接一个地讨论有关的主题。在非常简短地介绍了R和RStudio的安装和运行之后,它直接进入文本分析主题和技术,并综合讨论了R编程技巧。“快速和中肯”的特点旨在使“技术有用和立即回报”的读者(第vii页)。此外,书中使用的语言轻松幽默。作者更像是一个讲故事的人,而不是一个严肃的教授,总是用类比和隐喻来解释看似单调的技术零碎,这使得这本书显然是一个机智和愉快的阅读。本书分为三个部分。第一部分“微观分析”包括前五章。它主要涉及词频和分布,并介绍了R语言的大部分基本编程技术。第二部分“Mesoanalysis”由以下五章组成,探讨了更高层次的主题,如词汇多样性,hapax丰富性,KWIC和XML解析。第三部分“宏观分析”包含后三部分
Research in the area of literature is largely qualitatively based and people in the area may sometimes challenge the significance of using quantitative and computational methods in literature study. The author in the Preface of the book under review addressed the issue by arguing that results of the computational approaches could add complementary evidence to those of the “speculative methods” and might come up with new findings (p. viii). The book is one of the Quantitative Methods in the Humanities and Social Sciences Series. It aims to provide literary, or more generally humanist, researchers quantitative and computational methods of studying the texts (p. vii). Along with those methods, the book discusses the basics of R, a free but powerful language and tool for text analysis and statistics. Since most of its target audience may have no experience of programming, the book assumes its readers have no background knowledge of programming. With that said, the book is different from other excellent books of similar purposes such as Baayen (2008), Gries (2009) and Gries (2013) in its structure. It does not start with a brief introduction to the basic data types of the programming language and then discusses the topics concerned one after another. After a very brief introduction to the installation and running of R and RStudio, it goes straight to text analysis topics and techniques with integrated discussions on R programming skills. The “fast and tothe-point” feature aims to make “the technical useful and immediately rewarding” to the readers (p. vii). In addition, the language used in the book is easy and humorous. The author is more of a storyteller than a serious professor and always explains the seemingly monotonous technical odds and ends with analogies and metaphors, which makes the book an obviously witty and enjoyable read. The book is divided into three parts. The first part “Microanalysis” contains the first five chapters. It mainly deals with word frequency and distribution and introduces most of the basic programming techniques of the R language. The second part “Mesoanalysis”, which is composed of the next five chapters, explores topics of a higher level, such as lexical variety, hapax richness, KWIC and XML parsing. The third part “Macroanalysis” contains the last three