Collision between biological process and statistical analysis revealed by mean-centering.

Collision between biological process and statistical analysis revealed by mean-centering.
复制标题

均值中心化揭示了生物过程和统计分析之间的冲突。

DOI:
10.1111/1365-2656.13360
复制
发表时间:
2020
期刊:
The Journal of animal ecology
影响因子:
--
通讯作者:
H. Schielzeth
H. Schielzeth
中科院分区:
--
文献类型:
--
作者:
D. Westneat;Y. Araya;Hassen Allegue;B. Class;N. Dingemanse;N. Dochtermann;L. Garamszegi;Julien G. A. Martin;Shinichi Nakagawa;D. Réale;H. Schielzeth

文献摘要

参考文献

被引文献

相似文献

动物生态学家经常收集分层结构的数据,并用线性混合效应模型分析这些数据。当协变量的效应量在多个水平上变化时,会出现特定的并发症(例如,在受试者之间)。在这种情况下,受试者内协变量的均值中心化提供了一种有用的方法,但并非没有问题。统计模型代表了关于潜在生物过程的假设。聚类内的均值居中假设较低水平的响应(例如受试者内)取决于受试者均值(相对)的偏差,而不是协变量的绝对值。这在生物学上可能是现实的,也可能不是。我们表明,产生的性质(即,生物学)过程和统计分析的形式对生物学家产生了重大的概念和操作挑战。我们探讨了不匹配的后果,通过模拟数据与三个响应生成过程不同的协变量和响应之间的相关性的来源。然后用三个不同的分析方程对这些数据进行分析。我们询问了不同的分析方程如何稳健地估计感兴趣的关键参数,以及在哪些情况下会出现偏差。生成方程和解析方程之间的不匹配为估计关键参数带来了一些棘手的问题。最广泛的错误估计参数是反应的受试者间方差。我们发现,没有一个单一的分析方程是强大的估计所有方程产生的所有参数。重要的是,即使响应生成方程和分析方程在数学上匹配,当在协变量范围内进行采样时,某些参数的偏倚也会出现。我们的研究结果对我们如何收集和分析数据具有普遍意义。它们还更普遍地提醒我们,数据统计分析的结论是以一个假设为条件的,有时是隐含的,对于产生我们测量的属性的过程。我们讨论的策略,真实的数据分析在面对潜在的生物过程的不确定性。
Animal ecologists often collect hierarchically-structured data and analyze these with linear mixed-effects models. Specific complications arise when the effect sizes of covariates vary on multiple levels (e.g., within vs among subjects). Mean-centering of covariates within subjects offers a useful approach in such situations, but is not without problems. A statistical model represents a hypothesis about the underlying biological process. Mean-centering within clusters assumes that the lower level responses (e.g. within subjects) depend on the deviation from the subject mean (relative) rather than on absolute values of the covariate. This may or may not be biologically realistic. We show that mismatch between the nature of the generating (i.e., biological) process and the form of the statistical analysis produce major conceptual and operational challenges for empiricists. We explored the consequences of mismatches by simulating data with three response-generating processes differing in the source of correlation between a covariate and the response. These data were then analyzed by three different analysis equations. We asked how robustly different analysis equations estimate key parameters of interest and under which circumstances biases arise. Mismatches between generating and analytical equations created several intractable problems for estimating key parameters. The most widely misestimated parameter was the among-subject variance in response. We found that no single analysis equation was robust in estimating all parameters generated by all equations. Importantly, even when response-generating and analysis equations matched mathematically, bias in some parameters arose when sampling across the range of the covariate was limited. Our results have general implications for how we collect and analyze data. They also remind us more generally that conclusions from statistical analysis of data are conditional on a hypothesis, sometimes implicit, for the process(es) that generated the attributes we measure. We discuss strategies for real data analysis in face of uncertainty about the underlying biological process.
DOI: 10.3389/fevo.2017.00092
发表时间: 2017
影响因子: 3
作者:
P. Sprau;N. Dingemanse
通讯作者: P. Sprau;N. Dingemanse
DOI: 10.1111/2041-210x.12281
发表时间: 2015-01-01
影响因子: 6.6
作者:
Cleasby, Ian R.;Nakagawa, Shinichi;Schielzeth, Holger
通讯作者: Schielzeth, Holger
DOI: 10.1146/annurev.psych.093008.100356
发表时间: 2011
影响因子: 24.8
作者:
Curran PJ;Bauer DJ
通讯作者: Bauer DJ
DOI: 10.1037/a0012869
发表时间: 2008-09-01
影响因子: 7
作者:
Luedtke, Oliver;Marsh, Herbert W.;Muthen, Bengt
通讯作者: Muthen, Bengt