ProtAnt A tool for analysing the prototypicality of texts

ProtAnt A tool for analysing the prototypicality of texts
复制标题

ProtAnt 分析文本原型的工具

DOI:
10.1075/ijcl.20.3.01ant
复制
发表时间:
2015
影响因子:
1
通讯作者:
Anthony L
Anthony L
中科院分区:
人文科学4区
文献类型:
--
作者:
Anthony L

文献摘要

相似文献

基于语料库的研究人员和传统的定性研究人员,例如对批判性话语分析感兴趣的研究人员,经常需要选择原型文本进行细读,这些文本包括在更大的语料库中存在的感兴趣的语言特征。这一选拔程序的传统方法在很大程度上是临时的。在本文中,我们提供了一种更有原则的方法,根据文本包含的关键字数量对文本进行排名,以选择文本进行细读。为了便于分析,我们开发了一个名为ProtAnt的多平台免费软件工具,它可以分析文本,根据统计显著性和效应大小生成关键字排名列表,然后根据其中的关键字数量对文本进行排序。我们描述了各种实验,证明ProtAnt分析不仅在识别原型文本方面有效,而且在识别可能需要从目标语料库中删除的异常文本方面也有效。
Corpus-based researchers and traditional qualitative researchers, such as those interested in critical discourse analysis, are often required to select prototypical texts for close reading that include the language features of interest that are present in a much larger corpus. Traditional approaches to this selection procedure have been largely ad hoc. In this paper, we offer a more principled way of selecting texts for close reading based on a ranking of texts in terms of the number of keywords they contain. To facilitate this analysis, we have developed a multiplatform, freeware software tool called ProtAnt that analyses the texts, generates a ranked list of keywords based on statistical significance and effect size, and then orders the texts by the number of keywords in them. We describe various experiments that demonstrate the ProtAnt analysis is effective not only at identifying prototypical texts, but also identifying outlier texts that may need to be removed from a target corpus.