AN INVESTIGATION OF EFFECT OF MISCLASSIFICATION ON PROPERTIES OF X2-TESTS IN ANALYSIS OF CATEGORICAL DATA

AN INVESTIGATION OF EFFECT OF MISCLASSIFICATION ON PROPERTIES OF X2-TESTS IN ANALYSIS OF CATEGORICAL DATA
复制标题

DOI:
10.1093/biomet/52.1-2.95
复制
发表时间:
1965-01-01
期刊:
影响因子:
2.7
通讯作者:
ANDERSON, RL
ANDERSON, RL
中科院分区:
数学2区
文献类型:
--
作者:
MOTE, VL;ANDERSON, RL

文献摘要

被引文献

相似文献

“在更复杂的诊断中,临床医生意识到存在相当大的错误风险,这种风险可能会因所研究的疾病,诊断测试的可用性和存在以及其他因素而发生很大变化。Diamond & Lilienfeld(1962 a,B)和纽韦尔(1962)最近发表的文章都涉及流行病学研究中分类错误和诊断错误的影响。当有困难在适当的分类观察到各自的类别,出现以下问题:(i)我们如何考虑这些错误在“通常”的测试假设?(ii)如果忽略这些错误,显著性水平将受到怎样的影响?(iii)这些误差对检验的功效有什么影响?为了说明问题的本质,假设样本大小为4的样本取自一个二项总体,该总体假设具有给定特征(称为成功)的比例p。如果0或4个个体具有该特征,则拒绝p= I的零假设。假设没有错误分类和随机抽样,I类错误率为a= 0-125,该检验拒绝p= 0-25或0 - 75的零假设的功效为0-320,p= 0-1或0 - 9的功效为0-656。假设有25%的人被记录为没有这种特征,而没有这种特征的人中有25%被判断为有这种特征。在这些情况下,测试的功效确定如下。设r为记录的成功次数,而实际上有s次成功。s次成功的概率(4次中)为g(slp)= C(4,s)ps(l _p)4-s(s= 0,1,.其中C(4,s)是组合符号。设f(r)为如果有s次成功,则会有r次成功的概率。则(i)s= 0:fo(r)= C(4r)34-r/44. (ii)s= 1:从三次实际失败中记录的r1成功和从一次实际成功中记录的r8成功的卷积给出f1(r)= 33-r [C(3,r)+ 32 C(3,r-1)]/44。t C(n,m)= 0 if m < O or m>n.
'In more complex diagnoses, the clinician realizes that there is a considerable risk of error, a risk that may vary a great deal depending on the disease under study, the availability and existence of diagnostic tests, and other factors.'Recent articles by Diamond & Lilienfeld (1962 a, b) and by Newell (1962) are concerned with the effects of both classification and diagnosis errors in epidemiological studies. When there is difficulty in the proper classification of observations into the respective categories, the following questions arise:(i) How do we take these errors into account in the'usual'tests of hypotheses?(ii) If these errors are ignored, how will the level of significance be affected?(iii) What will be the effect of these errors on the power of the test? In order to illustrate the nature of the problem, suppose that samples of size 4 are taken from a binomial population assumed to have a proportion p of a given characteristic (called a success). If 0 or 4 individuals have this characteristic the null hypothesis of p= I is rejected. Assuming no misclassification and random sampling, the type I error rate is a= 0-125 and the power of this test to reject the null hypothesis for p= 0-25 or 0 75 is 0-320 and for p= 0-1 or 0 9 is 0-656. Suppose 25% of those with the characteristic are recorded as not having it and 25% of those without it are judged as having this characteristic. The power of the test under these circumstances is determined as follows. Let r be the number of recorded successes when there actually were s successes. The probability of s successes (out of 4) is g (slp)= C (4, s) ps (l _p) 4-s (s= 0, 1,... 4), where C (4, s) is the combinatorial symbol. Let f (r) be the probability that if there are s successes, there will be r recorded successes. Then (i) s= 0: fo (r)= C (4 r) 34-r/44.(ii) s= 1: the convolution of r1 recorded successes from the three actual failures and r8 recorded successes from the one actual success givest f1 (r)= 33-r [C (3, r)+ 32 C (3, r-1)]/44. t C (n, m)= Oif m< O or m> n.