z -squared: The Origin and Application of

z -squared: The Origin and Application of
复制标题

z 平方:起源和应用

DOI:
10.1080/09296174.2013.830554
复制
发表时间:
2013
影响因子:
1.4
通讯作者:
Wallis S
Wallis S
中科院分区:
人文科学4区
文献类型:
--
作者:
Wallis S

文献摘要

相似文献

在语言学研究中,经常使用一套统计检验方法,即列联检验,其中χ 2检验是最著名的例子。权变测试比较离散分布,即数据分为两个或两个以上的备选类别,如说话者的备选语言选择或不同的实验条件。这些测试非常普遍,是每个语言学研究者武器库的一部分。然而,这些测试的数学基础很少在文献中以一种平易近人的方式讨论,结果是许多研究人员可能不适当地应用测试,看不到测试特定问题的可能性,或者得出不合理的结论。列联检验也与置信区间的构造密切相关,置信区间对于绘制实验观测的确定性是非常有用和揭示性的方法。本文件的结构如下。本文介绍了最简单的χ 2检验的基本原理,即2 × 1拟合优度检验,并将其与单观察比例p的z检验和p的Wilson评分置信区间相联系。然后,我们展示了如何从两个观测值sp1和p2推导出独立性(同质性)的2 × 2检验,并解释了何时应该使用每种检验。我们还简要介绍了纽科姆-威尔逊检验,理想情况下,对于从两个独立群体(例如两个子语料库)中获得的观察结果,应优先使用该检验而不是χ检验。然后,我们转向更大的表的测试,通常称为r × c测试,它具有多个自由度,因此可能包含多个趋势,并讨论其分析策略。最后,我们简要地转向区分测试结果的问题。我们引入了效果大小的概念(也称为“关联的度量”),最后解释了我们如何进行纵向可分性检验来区分两组结果。
A set of statistical tests termedcontingency tests, of which χ2is the most well-known example, are commonly employed in linguistics research. Contingency tests compare discrete distributions, that is, data divided into two or more alternative categories, such as alternative linguistic choices of a speaker or different experimental conditions. These tests are highly ubiquitous, and are part of every linguistics researcher’s arsenal. However, the mathematical underpinnings of these tests are rarely discussed in the literature in an approachable way, with the result that many researchers may apply tests inappropriately, fail to see the possibility of testing particular questions, or draw unsound conclusions. Contingency tests are also closely related to the construction ofconfidence intervals, which are highly useful and revealing methods for plotting the certainty of experimental observations. This paper is organized in the following way. The foundations of the simplest type of χ2test, the 2 × 1 goodness of fit test, is introduced and related to theztest for a single observed proportionpand the Wilson score confidence interval aboutp. We then show how the 2 × 2 test for independence (homogeneity) is derived from two observationsp1andp2and explain when each test should be used. We also briefly introduce the Newcombe-Wilson test, which ideally should be used in preference to the χ test for observations drawn from two independent populations (such as two sub-corpora). We then turn to tests for larger tables, generally termedr×ctests, which have multiple degrees of freedom and therefore may encompass multiple trends, and discuss strategies for their analysis. Finally, we turn briefly to the question of differentiating test results. We introduce the concept ofeffect size(also termed “measures of association”) and finally explain how we may perform statisticalseparability teststo distinguish between two sets of results.