Repeatability for Gaussian and non-Gaussian data: a practical guide for biologists

Repeatability for Gaussian and non-Gaussian data: a practical guide for biologists
复制标题

DOI:
10.1111/j.1469-185x.2010.00141.x
复制
发表时间:
2010-11-01
期刊:
影响因子:
10
通讯作者:
Schielzeth, Holger
Schielzeth, Holger
中科院分区:
生物学1区
文献类型:
--
作者:
Nakagawa, Shinichi;Schielzeth, Holger

文献摘要

被引文献

相似文献

重复性(更准确地说,重复性的常见衡量标准是类内相关系数,ICC)是量化测量准确性和表型稳定性的重要指标。表型变异的比例可以归因于受试者之间(或组之间)的变异。因此,表型变异的不可重复部分是测量误差和表型灵活性的总和。有几种方法可以评估高斯数据的可重复性,但对于如何计算非高斯数据(如二进制、比例和计数数据)的可重复性,没有正式的协议。除了点估计之外,无论数据类型如何,都需要适当的不确定度估计(标准误差和可信区间)和重复性估计的统计学意义。我们回顾了计算高斯型和非高斯型数据的重复性和相关统计量的方法。对于高斯数据,我们提出了三种常用的方法来估计重复性:基于相关性的方法、基于方差分析(ANOVA)的方法和基于线性混合效应模型(LMM)的方法,而对于非高斯数据,我们专注于广义线性混合效应模型(GLMM),它们允许在原始和潜在尺度上估计重复性。我们还讨论了一些计算标准误差、可信区间和统计显著性的方法;最准确和最推荐的方法是参数自举、随机化测试和贝叶斯方法。我们提倡使用基于LMM和GLMM的方法,主要是因为可以很容易地控制混杂变量。此外,我们比较了两种类型的可重复性(普通可重复性和外推可重复性)与狭义遗传性的关系。这篇综述为生物学家计算高斯和非高斯数据的重复性和遗传性提供了指南和建议的集合。
Repeatability (more precisely the common measure of repeatability, the intra-class correlation coefficient, ICC) is an important index for quantifying the accuracy of measurements and the constancy of phenotypes. It is the proportion of phenotypic variation that can be attributed to between-subject (or between-group) variation. As a consequence, the non-repeatable fraction of phenotypic variation is the sum of measurement error and phenotypic flexibility. There are several ways to estimate repeatability for Gaussian data, but there are no formal agreements on how repeatability should be calculated for non-Gaussian data (e.g. binary, proportion and count data). In addition to point estimates, appropriate uncertainty estimates (standard errors and confidence intervals) and statistical significance for repeatability estimates are required regardless of the types of data. We review the methods for calculating repeatability and the associated statistics for Gaussian and non-Gaussian data. For Gaussian data, we present three common approaches for estimating repeatability: correlation-based, analysis of variance (ANOVA)-based and linear mixed-effects model (LMM)-based methods, while for non-Gaussian data, we focus on generalised linear mixed-effects models (GLMM) that allow the estimation of repeatability on the original and on the underlying latent scale. We also address a number of methods for calculating standard errors, confidence intervals and statistical significance; the most accurate and recommended methods are parametric bootstrapping, randomisation tests and Bayesian approaches. We advocate the use of LMM- and GLMM-based approaches mainly because of the ease with which confounding variables can be controlled for. Furthermore, we compare two types of repeatability (ordinary repeatability and extrapolated repeatability) in relation to narrow-sense heritability. This review serves as a collection of guidelines and recommendations for biologists to calculate repeatability and heritability from both Gaussian and non-Gaussian data.