Reliability: What type, please!

Reliability: What type, please!
复制标题

可靠性:请问是什么类型!

DOI:
10.1016/j.jshs.2012.11.001
复制
发表时间:
2013
影响因子:
11.7
通讯作者:
Weimo Zhu
Weimo Zhu
中科院分区:
医学1区
文献类型:
--
作者:
Weimo Zhu

文献摘要

被引文献

相似文献

正如我们在第一节研究方法课上学到的,效度和信度是任何测试、测量或评估中最重要的两个品质。与效度相比,信度实际上更重要,因为没有信度就没有效度。由于可靠性如此重要,今天几乎所有的研究期刊都有一些与可靠性相关的文章。不幸的是,许多这类文章都没有报告关于可靠性的一个重要信息——它的类型。此外,如果报告的可靠性类型,通常不支持其研究设计。为了充分理解为什么报告可靠性的类型和相关的研究设计是重要的,简要回顾可靠性的定义、可靠性的类型以及它们与错误的关系可能会有所帮助。可靠性通常被定义为“重复测试过程时测量结果的一致性”。假设一个应试者做了一次测试,被测的能力或潜在特征没有变化;然后假设同样的测试再次给同一个测试者。人们会期望这两个试验的分数应该非常相似。否则,测试将是不可靠的。根据经典测试理论2a,如果我们对一个考生进行多次测试,这个人的测试分数,即观察分数,将不会总是相同的。如果我们在频率分布中绘制分数,那么假设没有学习或疲劳效应,这个分布应该看起来像一个正态分布,大多数分数接近分数分布的中心(平均值),只有少数分数非常大或非常小(图1)。在这种情况下,平均值代表了测试者的能力或内在特征的水平,根据所采用的测试理论,它被称为“真实分数”,“宇宙分数”或“能力/特征”。观察到的分数与真实分数之间的距离通常被称为“误差”,它可以代表被测量能力的自然变化,也可以由某种系统误差引起。因此,任何观察到的分数在概念上都可以被认为包含两个部分:真实分数加上错误分数。当误差为零时,观察到的分数(图1中的x1)将等于真实分数。真实的分数在现实生活中是未知的,但可以通过确定测量误差并从获得的分数中减去它来估计。观察到的分数x2在真实分数的正侧误差略大,而观察到的分数x3的误差要大得多,但在负侧。因此,观察分数、真实分数和误差之间的关系可以概括为:观察分数(X)=真实分数(T)+误差(E)。
Validity and reliability, as we all learned in our first research methods class, are two of the most important qualities of any test, measurement or assessment. When compared with validity, reliability is actually more important since without it, there would be no validity. Since reliability is so important, almost all research journals today have some articles related to reliability. Unfortunately, many of these articles fail to report one of the important pieces of information regarding reliability–its type. In addition, if the type of reliability is reported, it is often not supported by its study design. To fully understand why reporting the type of reliability and the related study design is important, a short review on the definition of reliability, its types, and their relationship with errors may be helpful.Reliability is popularly defined as “the consistency of measurements when the testing procedure is repeated”. 1 Assume that a test taker did a test once and there is no change in the ability or underlying trait being measured; then suppose that the same test was administered again to that same test taker. One would expect the scores from these two trials should be quite similar. If not, the test would be unreliable. According to classical testing theory 2 a if we administer one test many times to a test taker, this person's test scores, known as the observed scores, will not be the same all the time. If we plot the scores in a frequency distribution, then assuming that there is no learning or fatigue effect, this distribution should look like a normal distribution with most scores close to the center (mean) of the score distribution, with a few very large or very small scores (Fig. 1). The mean in this case represents the level of ability or intrinsic traits of the test taker, which is known as the “true score”,“universe score”, or “ability/trait” depending on the testing theory employed. The distance between an observed score and the true score is often called “error”, which could represent natural variations in the ability being measured or may be caused by some sort of systematic error. Thus, any observed score can be conceptually considered to have two parts: a true score plus an error. When the error is zero, the observed score (X 1 in Fig. 1) will be equal to the true score. A true score is unknown in real life, but it can be estimated by determining the measurement error and subtracting it from the obtained score. The observed score X 2 has a slightly larger error on the positive side of the true score, whereas the observed score X 3 has a much larger error, but on the negative side. The relationship among the observed score, true score, and error can therefore be summarized as: observed score (X)= true score (T)+ error (E).