Best Practices for Binary and Ordinal Data Analyses.

Best Practices for Binary and Ordinal Data Analyses.
复制标题

DOI:
10.1007/s10519-020-10031-x
复制
发表时间:
2021-05
期刊:
影响因子:
2.6
通讯作者:
Neale MC
Neale MC
中科院分区:
医学3区
文献类型:
--
作者:
Verhulst B;Neale MC

文献摘要

参考文献

相似文献

许多人类特征、状态和障碍的测量都是从问卷上的一组项目开始的。这些问题的回答格式通常是简单的二进制(例如,是/否)或被命令(例如,高、中或低)。在数据分析过程中,这些项目经常被求和或用于估计因子得分。在临床应用中,此类评估在一般人群中通常呈非正态分布,因为许多受访者未受影响,因此无症状。因此,在许多情况下,这些措施违反了后续分析所需的统计假设。为了减少非正态性和准连续评估的影响,变量经常被重新编码为二元(受影响-未受影响)或有序(轻度-中度-重度)诊断。因此,有序数据在多个层次的分析中提出了挑战。将连续变量分类为有序类别通常会导致统计功效的损失,这表示数据分析师有动机假设数据是正态分布的,即使它们不是。尽管先前的时代精神学家认为,例如,具有超过10个有序类别的变量可以被认为是连续的,并且被分析,就好像它们是连续的一样,我们通过模拟研究表明,通常情况下并非如此。特别是,使用皮尔逊积矩相关性,而不是最大似然估计的多区域相关性的偏差估计的相关性向零。当多个观察结果落入单个观察类别(例如得分为零)时,这种偏差尤其严重。相比之下,用最大似然法估计有序相关性不会产生估计偏差,尽管标准误(适当地)更大。我们还说明了优势比如何严重依赖于受影响的个人在人口中的比例或患病率,因此是次优的研究,需要比较的关联指标。最后,我们将这些分析扩展到经典的双胞胎模型,并证明将二进制数据视为连续的会低估遗传和共同的环境方差分量,高估独特的环境(残差)方差。这些偏见随着流行率的下降而增加。虽然对有序数据进行适当的建模可能会更加计算密集和耗时,但如果不这样做,可能会产生有偏的相关性和有偏的参数估计。
The measurement of many human traits, states, and disorders begins with a set of items on a questionnaire. The response format for these questions is often simply binary (e.g., yes/no) or ordered (e.g., high, medium or low). During data analysis, these items are frequently summed or used to estimate factor scores. In clinical applications, such assessments are often non-normally distributed in the general population because many respondents are unaffected, and therefore asymptomatic. As a result, in many cases these measures violate the statistical assumptions required for subsequent analyses. To reduce the influence of the non-normality and quasi-continuous assessment, variables are frequently recoded into binary (affected-unaffected) or ordinal (mild-moderate-severe) diagnoses. Ordinal data therefore present challenges at multiple levels of analysis. Categorizing continuous variables into ordered categories typically results in a loss of statistical power, which represents an incentive to the data analyst to assume that the data are normally distributed, even when they are not. Despite prior zeitgeists suggesting that, e.g., variables with more than 10 ordered categories may be regarded as continuous and analyzed as if they were, we show via simulation studies that this is not generally the case. In particular, using Pearson product-moment correlations instead of maximum likelihood estimates of polychoric correlations biases the estimated correlations towards zero. This bias is especially severe when a plurality of the observations fall into a single observed category, such as a score of zero. By contrast, estimating the ordinal correlation by maximum likelihood yields no estimation bias, although standard errors are (appropriately) larger. We also illustrate how odds ratios depend critically on the proportion or prevalence of affected individuals in the population, and therefore are sub-optimal for studies where comparisons of association metrics are needed. Finally, we extend these analyses to the classical twin model and demonstrate that treating binary data as continuous will underestimate genetic and common environmental variance components, and overestimate unique environment (residual) variance. These biases increase as prevalence declines. While modeling ordinal data appropriately may be more computationally intensive and time consuming, failing to do so will likely yield biased correlations and biased parameter estimates from modeling them.
DOI: 10.1016/j.cell.2017.05.038
发表时间: 2017-06-15
期刊: Cell
影响因子: 64.5
作者:
Boyle EA;Li YI;Pritchard JK
通讯作者: Pritchard JK
DOI: 10.1007/s11336-010-9200-6
发表时间: 2011-04-01
期刊: Psychometrika
影响因子: 3
作者:
Boker S;Neale M;Maes H;Wilde M;Spiegel M;Brick T;Spies J;Estabrook R;Kenny S;Bates T;Mehta P;Fox J
通讯作者: Fox J
DOI: 10.1007/bf02293801
发表时间: 1981-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
BOCK, RD;AITKIN, M
通讯作者: AITKIN, M
DOI: 10.1037/1082-989x.9.3.301
发表时间: 2004-09-01
影响因子: 7
作者:
Mehta, PD;Neale, MC;Flay, BR
通讯作者: Flay, BR
DOI: 10.1007/s10530-015-0970-8
发表时间: 2015-12-01
影响因子: 2.9
作者:
Bradie, Johanna;Pietrobon, Adam;Leung, Brian
通讯作者: Leung, Brian