The Forgotten Semantics of Regression Modeling in Geography

The Forgotten Semantics of Regression Modeling in Geography
复制标题

DOI:
10.1111/gean.12199
复制
发表时间:
2021-01-01
影响因子:
3.6
通讯作者:
Atkinson,Peter M.
Atkinson,Peter M.
中科院分区:
地球科学3区
文献类型:
--
作者:
Comber,Alexis John;Harris,Paul;Atkinson,Peter M.

文献摘要

相似文献

本文关注的是与空间数据统计分析相关的语义。它采取了最简单的情况下预测的变量作为一个函数的协变量(s)x,其中predictedy始终是一个近似的y,只有永远是一个函数的x,因此,继承了许多空间特性的x,并说明了几个核心问题,使用“合成”遥感和“真实的”土壤的案例研究。回归模型的输出,因此,predictedy的含义,是不同的,由于(1)数据的选择:规格的x(包括协变量),支持的x(测量尺度和粒度),测量的x和错误的x,和(2)选择的模型,包括其功能形式和模型识别的方法。其中一些问题比其他问题得到更广泛的认识。因此,该研究定义了回归预测和推断受数据和模型选择影响的多种方式。这篇文章邀请研究人员停下来考虑预测的语义,它通常只不过是协变量x的缩放版本,并认为忽略这一点是天真的。
This article is concerned with the semantics associated with the statistical analysis of spatial data. It takes the simplest case of the prediction of variableyas a function of covariate(s)x, in which predictedyis always an approximation ofyand only ever a function ofx, thus, inheriting many of the spatial characteristics ofx, and illustrates several core issues using “synthetic” remote sensing and “real” soils case studies. The outputs of regression models and, therefore, the meaning of predictedy, are shown to vary due to (1) choices about data: the specification ofx(which covariates to include), the support ofx(measurement scales and granularity), the measurement ofxand the error ofx, and (2) choices about the model including its functional form and the method of model identification. Some of these issues are more widely recognized than others. Thus, the study provides definition to the multiple ways in which regression prediction and inference are affected by data and model choices. The article invites researchers to pause and consider the semantic meaning of predictedy, which is often nothing more than a scaled version of covariate(s)x, and argues that it is naïve to ignore this.