Prediction of driving ability: Are we building valid models?

Prediction of driving ability: Are we building valid models?
复制标题

DOI:
10.1016/j.aap.2015.01.013
复制
发表时间:
2015-04
期刊:
Accident; analysis and prevention
影响因子:
--
通讯作者:
P. Hoggarth;Carrie R. H. Innes;J. Dalrymple-Alford;Richard D. Jones
P. Hoggarth;Carrie R. H. Innes;J. Dalrymple-Alford;Richard D. Jones
中科院分区:
其他
文献类型:
--
作者:
P. Hoggarth;Carrie R. H. Innes;J. Dalrymple-Alford;Richard D. Jones

文献摘要

被引文献

相似文献

利用非公路指标预测道路驾驶能力是驾驶研究的一个重要目标。大多数分类模型的主要目标是确定少数能够高精度预测驾驶能力的越野变量。不幸的是,分类模型经常过度拟合研究样本,导致预测精度膨胀,对相关人群的泛化能力差,从而导致有效性差。许多驾驶研究没有报告足够的细节来确定模型过度拟合的风险,很少报告任何验证技术,这对测试模型的可泛化性至关重要。在回顾文献后,我们采用回归建模背景下的最佳实践技术,使用中等样本量(n= 279)生成了一个模型。然后随机选择逐渐变小的样本量,我们表明,参与者与自变量的低比例可能导致过度拟合的模型和关于模型准确性的虚假结论。我们得出结论,通过遵循一些准则可以构建更稳定的模型。
The prediction of on-road driving ability using off-road measures is a key aim in driving research. The primary goal in most classification models is to determine a small number of off-road variables that predict driving ability with high accuracy. Unfortunately, classification models are often over-fitted to the study sample, leading to inflation of predictive accuracy, poor generalization to the relevant population and, thus, poor validity. Many driving studies do not report sufficient details to determine the risk of model over-fitting and few report any validation technique, which is critical to test the generalizability of a model. After reviewing the literature, we generated a model using a moderately large sample size (n= 279) employing best practice techniques in the context of regression modelling. By then randomly selecting progressively smaller sample sizes we show that a low ratio of participants to independent variables can result in over-fitted models and spurious conclusions regarding model accuracy. We conclude that more stable models can be constructed by following a few guidelines.