Text categories and where you can stick them : A crude formality index

Text categories and where you can stick them : A crude formality index
复制标题

文本类别以及可以粘贴它们的位置:粗略的正式索引

DOI:
10.1075/ijcl.2.2.04sig
复制
发表时间:
1997
影响因子:
1
通讯作者:
R. Sigley
R. Sigley
中科院分区:
人文科学4区
文献类型:
--
作者:
R. Sigley

文献摘要

被引文献

相似文献

本文应用主成分分析(PCA)来解决语言变异分析中已有语料文本类别的解释问题。该方法是通过构建一个索引的复杂的概念“正式”从PCA的一组高频基于单词的计数。第一个主要成分从这个分析作为一个广泛的正式指数;第二个主要成分被暂时确定为标记“具体的事实”与“抽象的讨论”。随后,从语料库中的文本类别定位在这些文本的维度,并选择类别的内部一致性进行评估,通过比较跨子类别的文本的分布。最后,本文对该方法的进一步发展和应用以及语料库在变异研究中的应用提出了建议。
This paper applies principal components analysis (PCA) to solve the problem of interpreting pre-existing corpus text categories for analysis of linguistic variation. The method is illustrated by constructing an index of the complex notion "formality " from PCA of a set of high-frequency wordform-based counts. The first principal component from this analysis acts as a broad formality index; a second principal component is tentatively identified as marking "concrete facts" versus "abstract discussion"'. Subsequently, text categories from the corpora are positioned on these textual dimensions, and selected categories are evaluated for internal consistency by comparing the distribution of texts across subcategories. Finally, suggestions are made concerning further developments and applications of the method used here, and its implications for the use of corpora in variation studies.