Exchangeability and regression models

Exchangeability and regression models
复制标题

可交换性和回归模型

DOI:
10.1093/acprof:oso/9780198566540.003.0005
复制
发表时间:
2004
期刊:
影响因子:
3.5
通讯作者:
P. McCullagh
P. McCullagh
中科院分区:
生物学3区
文献类型:
--
作者:
P. McCullagh

文献摘要

被引文献

相似文献

大卫·考克斯爵士的统计生涯以及他对随机过程理论和应用的终生兴趣始于羊毛工业的问题。将一根羊毛纱线牵伸到接近均匀宽度的问题并不是一个吉祥的起点,但 Cox (1949) 中使用了一系列来自平稳时间序列的令人印象深刻的时间和光谱方法来解决这个问题。他从平凡中提取基本原理的能力在他发现或构建了同名 Cox 过程(用于计算毛纱样品中的棉结)时变得显而易见(Cox,1955)。随后的应用包括水文学和远程依赖性(Davison 和 Cox 1989;Cox 1991)、降雨模型(Cox 和 Isham 1988;Rodriguez-Iturbe、Cox 和 Isham 1987、1988)以及传染病传播模型(Anderson、Cox 和 Hillier 1989)。在 20 世纪 50 年代末的某个时候,重点转向了依赖性统计模型,即响应变量依赖于已知解释变量或因素的方式(Cox,1958a)。他在这两个领域的贡献都非常有洞察力,Cox 过程是点过程的基本类别,并且 Cox 模型在生存分析中发挥着类似的作用。除此之外,我们还有 Box-Cox 变换、二元回归模型 (Cox 1958b) 和与农业田间试验相关的模型。这个简短的总结是对大卫爵士工作的粗略简化,但它通过介绍的方式符合我的目的,因为本章的主要目标是探索可交换性(随机过程的概念)与回归模型(其中观察到的过程由协变量调节)之间的关系。通常将随机过程的概念引入为随机变量 Y1、Y2 的集合。 。 .,通常是无限集,但不一定是有序序列。这意味着 U 是统计单位的索引集,并且对于每个有限子集 S = {u1, . 。 。 U 中的元素的值 Y (S) = ( Y (u1), ..., Y (un) ) 在 S 上的过程在 RS 上具有分布 PS 。本章强调概率分布而不是随机变量。因此,实值过程是对观测空间的概率分布的一致分配,使得在删除相关坐标的情况下,Rn 上的分布 Pn 是 Rn+1 上 Pn+1 的边缘分布。像 Rn 这样强调观察空间维度的符号并不完全令人满意,因为两个大小相等的样本不需要具有相同的分布,因此我们将采样单元上的实值函数集写为 RS 而不是 Rn。如果每个有限维分布是对称的,或者在坐标排列下不变,则称过程是可交换的。该定义表明,可交换性在依赖性统计模型中不起任何作用,在该模型中,由于协变量值的差异,分布明显不可交换。我认为
Sir David Cox’s statistical career and his lifelong interest in the theory and application of stochastic processes began with problems in the wool industry. The problem of drafting a strand of wool yarn to near uniform width is not an auspicious starting point, but an impressive array of temporal and spectral methods from stationary time series were brought to bear on the problem in Cox (1949). His ability to extract the fundamental from the mundane became evident in his discovery or construction of the eponymous Cox process in the counting of neps in a sample of wool yarn (Cox, 1955). Subsequent applications included hydrology and long-range dependence (Davison and Cox 1989; Cox 1991), models for rainfall (Cox and Isham 1988; Rodriguez-Iturbe, Cox and Isham 1987, 1988), and models for the spread of infectious diseases (Anderson, Cox and Hillier 1989). At some point in the late 1950s, the emphasis shifted to statistical models for dependence, the way in which a response variable depends on known explanatory variables or factors (Cox, 1958a). His contributions in both areas have been extraordinarily insightful, Cox processes being a fundamental class of point processes, and the Cox model playing a similar role in survival analysis. In addition to these, we have the Box-Cox transformation, binary regression models (Cox 1958b) and models relevant to agricultural field trials. This brief summary is a gross simplification of Sir David’s work, but it suits my purpose by way of introduction because the chief goal of this chapter is to explore the relation between exchangeability, a concept from stochastic processes, and regression models in which the observed process is modulated by a covariate. It is usual to introduce the notion of a stochastic process as a collection of random variables, Y1, Y2, . . ., usually an infinite set though not necessarily an ordered sequence. What this means is that U is an index set of statistical units, and for each finite subset S = {u1, . . . , un} of elements in U , the value Y (S) = ( Y (u1), . . . , Y (un) ) of the process on S has distribution PS on RS . This chapter emphasizes probability distributions rather than random variables. A real-valued process is thus a consistent assignment of probability distributions to observation spaces such that the distribution Pn on Rn is the marginal distribution of Pn+1 on Rn+1 under deletion of the relevant coordinate. A notation such as Rn that emphasizes the dimension of the observation space is not entirely satisfactory because two samples of equal size need not have the same distribution, so we write RS rather than Rn for the set of real-valued functions on the sampled units. A process is said to be exchangeable if each finite-dimensional distribution is symmetric, or invariant under coordinate permutation. The definition suggests that exchangeability can have no role in statistical models for dependence, in which the distributions are overtly non-exchangeable on account of differences in covariate values. I argue that