Statistics notes - The cost of dichotomising continuous variables

Statistics notes - The cost of dichotomising continuous variables
复制标题

DOI:
10.1136/bmj.332.7549.1080
复制
发表时间:
2006-05-06
影响因子:
--
通讯作者:
Royston, P
Royston, P
中科院分区:
医学1区
文献类型:
--
作者:
Altman, DG;Royston, P

文献摘要

被引文献

相似文献

连续变量的测量是在医学的所有分支中进行的,有助于患者的诊断和治疗。在临床实践中,根据连续变量的值,将个体标记为具有或不具有某种属性,如“高血压”或“肥胖”或“高胆固醇”是有帮助的。对连续变量的分类在临床研究中也很常见,但在这里,这种简单性是以一定的成本获得的。尽管分组可能有助于数据的呈现,尤其是在表格中,但分类对于统计分析来说是不必要的,而且它有一些严重的缺陷。在这里,我们考虑将连续数据转换为两组(二分法)的影响,因为这是临床研究中最常见的方法。1强迫所有人分成两组有什么公认的好处?一种常见的论点是,它极大地简化了统计分析,并导致易于解释和展示结果。二元分裂--例如,在中位数--导致具有较高或较低测量值的个体分组的比较,在最简单的情况下导致t检验或χ2检验以及对分组之间差异的估计(及其可信区间)。然而,总体上没有充分的理由假设存在潜在的二分法,如果存在的话,也没有理由认为它应该处于中位数。2双眼分割法导致了几个问题。首先,大量信息丢失,因此检测变量和患者预后之间关系的统计能力降低。事实上,在中位数对变量进行二分法所减少的电量与丢弃三分之一的数据所减少的电量相同。2、当研究研究已经趋于太小的时候,刻意丢弃数据肯定是不可取的。二分法还可能增加阳性结果为假阳性的风险。4其次,人们可能严重低估了不同组之间结果的差异程度,例如某些事件的风险,并且每个组内可能包含相当大的变异性。接近切点但在切点两边的个体被描述为非常不同,而不是非常相似。第三,使用两组数据隐藏了变量与结果之间的任何非线性关系。想必,许多将其一分为二的人并没有意识到其中的影响。如果使用二分法,切入点应该在哪里?对于一些变量,有可识别的分界点,例如以体重指数为基础定义“超重”的25公斤/平方米。对于一些变量,如年龄,通常会取一个整数,通常是5或10的倍数。可能会采用以前研究中使用的切点。在没有事先截止点的情况下,最常见的方法是取样本中位数。然而,使用样本中位数意味着在不同的研究中会使用不同的切点,因此它们的结果不容易进行比较,这严重阻碍了观察性研究的荟萃分析。然而,所有这些方法都是
Measurements of continuous variables are made in all branches of medicine, aiding in the diagnosis and treatment of patients. In clinical practice it is helpful to label individuals as having or not having an attribute, such as being “hypertensive” or “obese” or having” high cholesterol,” depending on the value of a continuous variable.Categorisation of continuous variables is also common in clinical research, but here such simplicity is gained at some cost. Though grouping may help data presentation, notably in tables, categorisation is unnecessary for statistical analysis and it has some serious drawbacks. Here we consider the impact of converting continuous data to two groups (dichotomising), as this is the most common approach in clinical research. 1 What are the perceived advantages of forcing all individuals into two groups? A common argument is that it greatly simplifies the statistical analysis and leads to easy interpretation and presentation of results. A binary split—for example, at the median—leads to a comparison of groups of individuals with high or low values of the measurement, leading in the simplest case to a t test or χ2 test and an estimate of the difference between the groups (with its confidence interval). There is, however, no good reason in general to suppose that there is an underlying dichotomy, and if one exists there is no reason why it should be at the median. 2 Dichotomising leads to several problems. Firstly, much information is lost, so the statistical power to detect a relation between the variable and patient outcome is reduced. Indeed, dichotomising a variable at the median reduces power by the same amount as would discarding a third of the data. 2 3 Deliberately discarding data is surely inadvisable when research studies already tend to be too small. Dichotomisation may also increase the risk of a positive result being a false positive. 4 Secondly, one may seriously underestimate the extent of variation in outcome between groups, such as the risk of some event, and considerable variability may be subsumed within each group. Individuals close to but on opposite sides of the cutpoint are characterised as being very different rather than very similar. Thirdly, using two groups conceals any non-linearity in the relation between the variable and outcome. Presumably, many who dichotomise are unaware of the implications. If dichotomisation is used where should the cutpoint be? For a few variables there are recognised cutpoints, such as> 25 kg/m2 to define “overweight” based on body mass index. For some variables, such as age, it is usual to take a round number, usually a multiple of five or 10. The cutpoint used in previous studies may be adopted. In the absence of a prior cutpoint the most common approach is to take the sample median. However, using the sample median implies that various cutpoints will be used in different studies so that their results cannot easily be compared, seriously hampering meta-analysis of observational studies. 5 Nevertheless, all these approaches are