Rs for Correlated Data: Phylogenetic Models, LMMs, and GLMMs

Rs for Correlated Data: Phylogenetic Models, LMMs, and GLMMs
复制标题

DOI:
10.1093/sysbio/syy060
复制
发表时间:
2019-03-01
期刊:
影响因子:
6.5
通讯作者:
Ives, Anthony R.
Ives, Anthony R.
中科院分区:
生物学1区
文献类型:
--
作者:
Ives, Anthony R.

文献摘要

被引文献

相似文献

许多研究人员希望报告一个模型来衡量模型解释的方差。当模型包含数据之间的相关性时,例如系统发育模型和混合模型,定义面临两个概念问题。 (i) 当模型包含协方差时,尚不清楚如何测量由预测(独立)变量解释的方差。 (ii) 研究人员可能希望通过提出以下问题来包括由协方差解释的方差:有多少数据是由系统发育解释的?在这里,我研究了系统发育模型和混合模型的三个模型。是普通最小二乘法的扩展,通过模型估计的方差和协方差对残差进行加权;它与 Nakakawa 和 Schielzeth 提出的方法密切相关(2013 年。从广义线性混合效应模型中获取 R2 的通用且简单的方法。Methods Ecol. Evol. 4:133-142)。基于根据拟合模型预测每个残差并计算观察值和预测值之间的方差。基于拟合模型的可能性,因此反映了模型包含的信息量。这三个 s 被表述为部分 s,使得可以比较预测变量和方差分量(系统发育信号和随机效应)对模型拟合的贡献。由于部分 s 将完整模型与没有完整模型组件的简化模型进行比较,因此它们与划分方差的加性组件的边际 s 不同。我使用连续和二元响应数据的模拟(系统发育广义最小二乘法和系统发育逻辑回归)评估了系统发育模型的 s 属性。由于 s 广泛设计用于相关数据的任何模型,因此我还比较了线性混合模型和广义线性混合模型的 s。 、 、 和 在描述由模型的不同组成部分解释的方差方面都具有相似的性能。然而,对于模型解释数据中的方差有多大的问题给出了最直接的答案。最适合比较适合不同数据集的模型,因为它不依赖于样本大小。并且最适合评估应用于相同数据的同一模型中不同组件的重要性,因为它与统计显着性检验最密切相关。
Many researchers want to report an to measure the variance explained by a model. When the model includes correlation among data, such as phylogenetic models and mixed models, defining an faces two conceptual problems. (i) It is unclear how to measure the variance explained by predictor (independent) variables when the model contains covariances. (ii) Researchers may want the to include the variance explained by the covariances by asking questions such as How much of the data is explained by phylogeny? Here, I investigated three s for phylogenetic and mixed models. is an extension of the ordinary least-squares that weights residuals by variances and covariances estimated by the model; it is closely related to presented by Nakagawa and Schielzeth (2013. A general and simple method for obtaining R2 from generalized linear mixed-effects models. Methods Ecol. Evol. 4:133-142). is based on predicting each residual from the fitted model and computing the variance between observed and predicted values. is based on the likelihood of fitted models, and therefore, reflects the amount of information that the models contain. These three s are formulated as partial s, making it possible to compare the contributions of predictor variables and variance components (phylogenetic signal and random effects) to the fit of models. Because partial s compare a full model with a reduced model without components of the full model, they are distinct from marginal s that partition additive components of the variance. I assessed the properties of the s for phylogenetic models using simulations for continuous and binary response data (phylogenetic generalized least squares and phylogenetic logistic regression). Because the s are designed broadly for any model for correlated data, I also compared s for linear mixed models and generalized linear mixed models. , , and all have similar performance in describing the variance explained by different components of models. However, gives the most direct answer to the question of how much variance in the data is explained by a model. is most appropriate for comparing models fit to different data sets, because it does not depend on sample sizes. And is most appropriate to assess the importance of different components within the same model applied to the same data, because it is most closely associated with statistical significance tests.