Text Analysis with R for Students of Literature
Text Analysis with R for Students of Literature
复制标题
文学专业学生使用 R 进行文本分析
DOI:
--
复制
发表时间:
2016
影响因子:
1.4
通讯作者:
L. Lei
中科院分区:
文献类型:
--
作者:
L. Lei
Research in the area of literature is largely qualitatively based and people in the area may sometimes challenge the significance of using quantitative and computational methods in literature study. The author in the Preface of the book under review addressed the issue by arguing that results of the computational approaches could add complementary evidence to those of the “speculative methods” and might come up with new findings (p. viii). The book is one of the Quantitative Methods in the Humanities and Social Sciences Series. It aims to provide literary, or more generally humanist, researchers quantitative and computational methods of studying the texts (p. vii). Along with those methods, the book discusses the basics of R, a free but powerful language and tool for text analysis and statistics. Since most of its target audience may have no experience of programming, the book assumes its readers have no background knowledge of programming. With that said, the book is different from other excellent books of similar purposes such as Baayen (2008), Gries (2009) and Gries (2013) in its structure. It does not start with a brief introduction to the basic data types of the programming language and then discusses the topics concerned one after another. After a very brief introduction to the installation and running of R and RStudio, it goes straight to text analysis topics and techniques with integrated discussions on R programming skills. The “fast and tothe-point” feature aims to make “the technical useful and immediately rewarding” to the readers (p. vii). In addition, the language used in the book is easy and humorous. The author is more of a storyteller than a serious professor and always explains the seemingly monotonous technical odds and ends with analogies and metaphors, which makes the book an obviously witty and enjoyable read. The book is divided into three parts. The first part “Microanalysis” contains the first five chapters. It mainly deals with word frequency and distribution and introduces most of the basic programming techniques of the R language. The second part “Mesoanalysis”, which is composed of the next five chapters, explores topics of a higher level, such as lexical variety, hapax richness, KWIC and XML parsing. The third part “Macroanalysis” contains the last three