Rs for Correlated Data: Phylogenetic Models, LMMs, and GLMMs
Rs for Correlated Data: Phylogenetic Models, LMMs, and GLMMs
复制标题
DOI:
10.1093/sysbio/syy060
复制
发表时间:
2019-03-01
影响因子:
6.5
通讯作者:
Ives, Anthony R.
中科院分区:
文献类型:
--
作者:
Ives, Anthony R.
Many researchers want to report an to measure the variance explained by a model. When the model includes correlation among data, such as phylogenetic models and mixed models, defining an faces two conceptual problems. (i) It is unclear how to measure the variance explained by predictor (independent) variables when the model contains covariances. (ii) Researchers may want the to include the variance explained by the covariances by asking questions such as How much of the data is explained by phylogeny? Here, I investigated three s for phylogenetic and mixed models. is an extension of the ordinary least-squares that weights residuals by variances and covariances estimated by the model; it is closely related to presented by Nakagawa and Schielzeth (2013. A general and simple method for obtaining R2 from generalized linear mixed-effects models. Methods Ecol. Evol. 4:133-142). is based on predicting each residual from the fitted model and computing the variance between observed and predicted values. is based on the likelihood of fitted models, and therefore, reflects the amount of information that the models contain. These three s are formulated as partial s, making it possible to compare the contributions of predictor variables and variance components (phylogenetic signal and random effects) to the fit of models. Because partial s compare a full model with a reduced model without components of the full model, they are distinct from marginal s that partition additive components of the variance. I assessed the properties of the s for phylogenetic models using simulations for continuous and binary response data (phylogenetic generalized least squares and phylogenetic logistic regression). Because the s are designed broadly for any model for correlated data, I also compared s for linear mixed models and generalized linear mixed models. , , and all have similar performance in describing the variance explained by different components of models. However, gives the most direct answer to the question of how much variance in the data is explained by a model. is most appropriate for comparing models fit to different data sets, because it does not depend on sample sizes. And is most appropriate to assess the importance of different components within the same model applied to the same data, because it is most closely associated with statistical significance tests.