The relative importance of persons, items, subtests and languages to TOEFL test variance

The relative importance of persons, items, subtests and languages to TOEFL test variance
复制标题

人员、项目、子测试和语言对托福考试方差的相对重要性

DOI:
10.1177/026553229901600205
复制
发表时间:
1999
期刊:
影响因子:
4.1
通讯作者:
J. D. Brown
J. D. Brown
中科院分区:
人文科学2区
文献类型:
--
作者:
J. D. Brown

文献摘要

被引文献

相似文献

本项目的目的是探索不同人数、项目、分测验、语言及其各种相互作用对托福成绩可靠性(类似于经典理论的可靠性)的相对贡献。为此,提出了三个研究问题:(1)分布的特征是什么,以及整个测试及其子测试的经典理论信度估计有多高?(2)对于15种语言中的每一种,人、项目、子测验及其相互作用对测试方差的相对贡献是多少?(3)在所有15种语言中,人员、项目、子测试和语言以及它们之间的各种相互作用对测试方差的相对贡献是什么?该研究从托福通用数据集的24500名参与者中抽取了15000名考生,其中1000名来自15种不同的语言背景,该数据集本身是1991年5月托福全球管理的样本。该测试在正常操作条件下进行,包括所有三个子测试:(1)听力理解,(2)结构和书面表达,以及(3)词汇和阅读理解。分析包括描述性统计,经典理论的信度估计,并进行了一系列的概化研究,以隔离的方差分量,由于人,项目,子测试和语言,及其对测试的可靠性的影响。与以前的研究不同,这里的结果表明,当与其他重要的方差来源(人,项目和子测试)一起考虑时,语言差异仅占托福考试方差的很小一部分。这些结果应该证明是有用的测试开发人员和研究人员感兴趣的这些因素对测试设计的相对影响。
The purpose of this project was to explore the relative contributions to TOEFL score dependability (which is analogous to classical theory reliability) of various numbers of persons, items, subtests, languages and their various interactions. To these ends, three research questions were formulated: (1) What are the characteristics of the distributions, and how high are the classical theory reliability estimates for the whole test and its subtests? (2) For each of the 15 languages, what are the relative contributions to test variance of persons, items, subtests and their interactions? (3) Across all 15 languages, what are the relative contributions to test variance of persons, items, subtests and languages, as well as their various interactions? The study sampled 15 000 test takers, 1000 each from 15 different language backgrounds, from the total of 24 500 participants in the TOEFL generic data set which itself was a sample from the May 1991 worldwide administration of the TOEFL. The test was administered under normal operational conditions and included all three subtests: (1) Listening Comprehension, (2) Structure and Written Expression, and (3) Vocabulary and Reading Comprehension. The analyses included descriptive statistics, classical theory reliability estimates, and a series of generalizability studies conducted to isolate the variance components due to persons, items, subtests and languages, and their effects on the dependability of the test. Unlike previous research, the results here indicate that, when considered in concert with other important sources of variance (persons, items and subtests), language differences alone account for only a very small proportion of TOEFL test variance. These results should prove useful to test developers and researchers interested in the relative effects of such factors on test design.