RMSE is not enough: Guidelines to robust data-model comparisons for magnetospheric physics

RMSE is not enough: Guidelines to robust data-model comparisons for magnetospheric physics
复制标题

DOI:
10.1016/j.jastp.2021.105624
复制
发表时间:
2021-04-08
影响因子:
1.9
通讯作者:
Mukhopadhyay, Agnit
Mukhopadhyay, Agnit
中科院分区:
地球科学4区
文献类型:
--
作者:
Liemohn, Michael W.;Shane, Alexander D.;Mukhopadhyay, Agnit

文献摘要

被引文献

相似文献

磁层物理研究界在进行研究调查时使用广泛的定量数据模型比较方法(指标)。但通常情况下,任何特定研究都只会使用一两个指标,其中最常见的两个是皮尔逊相关系数和均方根误差 (RMSE)。由于指标旨在测试数据模型关系的特定方面,因此将比较限制为仅一两个指标会减少可以从分析中收集到的物理见解,从而限制了建模研究的可能结果。当应用多种类型的指标时,可以获得额外的物理见解。我们将指标分为两个主要组:1)拟合性能指标,通常基于数据模型价值差异; 2) 事件检测指标,它使用由指定阈值确定的数据和模型值的离散事件分类。除了这些组之外,根据度量评估的数据模型关系方面,还有几个主要的度量类别:1)准确性; 2)偏见; 3)精度; 4)协会; 5)和极端。另一类是技能,它是根据参考模型的性能来衡量这些指标中的任何一个。这些可以应用于数据或模型值的子集,称为可靠性和区分度评估。在磁层物理示例的背景下,我们讨论为特定研究选择指标的最佳实践。
The magnetospheric physics research community uses a broad array of quantitative data-model comparison methods (metrics) when conducting their research investigations. It is often the case, though, that any particular study will only use one or two metrics, with the two most common being Pearson correlation coefficient and root mean square error (RMSE). Because metrics are designed to test a specific aspect of the data-model relationship, limiting the comparison to only one or two metrics reduces the physical insights that can be gleaned from the analysis, restricting the possible findings from modeling studies. Additional physical insights can be obtained when many types of metrics are applied. We organize metrics into two primary groups: 1) fit performance metrics, often based on the data-model value difference; and 2) event detection metrics, which use a discrete event classification of data and model values determined by a specified threshold. In addition to these groups, there are several major categories of metrics based on the aspect of the data-model relationship that the metric assesses: 1) accuracy; 2) bias; 3) precision; 4) association; 5) and extremes. Another category is skill, which is a measure of any of these metrics against the performance of a reference model. These can be applied to a subset of either the data or the model values, known as reliability and discrimination assessments. In the context of magnetospheric physics examples, we discuss best practices for choosing metrics for particular studies.