Reliability: What type, please!
Reliability: What type, please!
复制标题
可靠性:请问是什么类型!
DOI:
10.1016/j.jshs.2012.11.001
复制
发表时间:
2013
影响因子:
11.7
通讯作者:
Weimo Zhu
中科院分区:
文献类型:
--
作者:
Weimo Zhu
Validity and reliability, as we all learned in our first research methods class, are two of the most important qualities of any test, measurement or assessment. When compared with validity, reliability is actually more important since without it, there would be no validity. Since reliability is so important, almost all research journals today have some articles related to reliability. Unfortunately, many of these articles fail to report one of the important pieces of information regarding reliability–its type. In addition, if the type of reliability is reported, it is often not supported by its study design. To fully understand why reporting the type of reliability and the related study design is important, a short review on the definition of reliability, its types, and their relationship with errors may be helpful.Reliability is popularly defined as “the consistency of measurements when the testing procedure is repeated”. 1 Assume that a test taker did a test once and there is no change in the ability or underlying trait being measured; then suppose that the same test was administered again to that same test taker. One would expect the scores from these two trials should be quite similar. If not, the test would be unreliable. According to classical testing theory 2 a if we administer one test many times to a test taker, this person's test scores, known as the observed scores, will not be the same all the time. If we plot the scores in a frequency distribution, then assuming that there is no learning or fatigue effect, this distribution should look like a normal distribution with most scores close to the center (mean) of the score distribution, with a few very large or very small scores (Fig. 1). The mean in this case represents the level of ability or intrinsic traits of the test taker, which is known as the “true score”,“universe score”, or “ability/trait” depending on the testing theory employed. The distance between an observed score and the true score is often called “error”, which could represent natural variations in the ability being measured or may be caused by some sort of systematic error. Thus, any observed score can be conceptually considered to have two parts: a true score plus an error. When the error is zero, the observed score (X 1 in Fig. 1) will be equal to the true score. A true score is unknown in real life, but it can be estimated by determining the measurement error and subtracting it from the obtained score. The observed score X 2 has a slightly larger error on the positive side of the true score, whereas the observed score X 3 has a much larger error, but on the negative side. The relationship among the observed score, true score, and error can therefore be summarized as: observed score (X)= true score (T)+ error (E).