Handling Skewed Data: A Comparison of Two Popular Methods

Handling Skewed Data: A Comparison of Two Popular Methods
复制标题

DOI:
10.3390/app10186247
复制
发表时间:
2020-09-01
影响因子:
2.7
通讯作者:
Kheirallah, Khalid A.
Kheirallah, Khalid A.
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Hammouri, Hanan M.;Sabo, Roy T.;Kheirallah, Khalid A.

文献摘要

被引文献

相似文献

生物医学和社会心理学研究的科学家需要一直处理扭曲的数据。在比较两组平均值的情况下,对数变换通常用作在利用两组检验之前将偏斜数据归一化的传统技术。另一种不假设正态性的方法是广义线性模型(GLM)结合适当的链接函数。在这项工作中,这两种技术进行了比较,使用蒙特卡罗模拟,每个由许多迭代,模拟两组偏斜的数据为三种不同的抽样分布:伽玛,指数和β。之后,这两种方法进行了比较的第一类错误率,功率率和估计的平均差异。我们的结论是,与对数转换的t检验有上级性能优于GLM方法的任何数据,是不正常的,并遵循β或γ分布。另外,对于指数分布的数据,GLM方法具有优于对数变换的t检验的上级性能。
Scientists in biomedical and psychosocial research need to deal with skewed data all the time. In the case of comparing means from two groups, the log transformation is commonly used as a traditional technique to normalize skewed data before utilizing the two-groupt-test. An alternative method that does not assume normality is the generalized linear model (GLM) combined with an appropriate link function. In this work, the two techniques are compared using Monte Carlo simulations; each consists of many iterations that simulate two groups of skewed data for three different sampling distributions: gamma, exponential, and beta. Afterward, both methods are compared regarding Type I error rates, power rates and the estimates of the mean differences. We conclude that thet-test with log transformation had superior performance over the GLM method for any data that are not normal and follow beta or gamma distributions. Alternatively, for exponentially distributed data, the GLM method had superior performance over thet-test with log transformation.