Three myths about risk thresholds for prediction models

Three myths about risk thresholds for prediction models
复制标题

DOI:
10.1186/s12916-019-1425-3
复制
发表时间:
2019-10-25
期刊:
影响因子:
9.3
通讯作者:
Van Calster, Ben
Van Calster, Ben
中科院分区:
医学1区
文献类型:
--
作者:
Wynants, Laure;van Smeden, Maarten;Van Calster, Ben

文献摘要

被引文献

相似文献

临床预测模型可用于根据患者当前的特征估计患者将来患某种疾病或经历某种事件的风险。定义适当的风险阈值以推荐干预是将风险预测模型引入临床应用的关键挑战;这种风险阈值通常以特定方式定义。这是有问题的,因为默认的假阳性和假阴性分类的成本可能在临床上不合理。例如,当选择使正确分类的患者比例最大化的风险阈值时,假阳性和假阴性被假设为同样昂贵。此外,小到中等的样本量可能会导致不稳定的最佳阈值,这需要特别谨慎的结果解释。我们讨论了三种常见的关于风险阈值的误解是如何导致患者不适当的风险分层的。首先,我们指出咨询和共同决策的背景下,连续的风险估计比风险分层更有用。其次,我们认为,阈值的选择应该反映风险分层后作出的决定的后果。第三,我们强调,通常没有普遍的最佳阈值,而是一个合理的风险阈值取决于临床背景。因此,我们建议在开发或验证预测模型时提供多个风险阈值的结果。结论牢记这三点可以避免干预措施的不适当分配(和不分配)。如果使用上下文相关的阈值,使用有区别的和校准良好的模型将产生更好的临床结果。
Background Clinical prediction models are useful in estimating a patient's risk of having a certain disease or experiencing an event in the future based on their current characteristics. Defining an appropriate risk threshold to recommend intervention is a key challenge in bringing a risk prediction model to clinical application; such risk thresholds are often defined in an ad hoc way. This is problematic because tacitly assumed costs of false positive and false negative classifications may not be clinically sensible. For example, when choosing the risk threshold that maximizes the proportion of patients correctly classified, false positives and false negatives are assumed equally costly. Furthermore, small to moderate sample sizes may lead to unstable optimal thresholds, which requires a particularly cautious interpretation of results. Main text We discuss how three common myths about risk thresholds often lead to inappropriate risk stratification of patients. First, we point out the contexts of counseling and shared decision-making in which a continuous risk estimate is more useful than risk stratification. Second, we argue that threshold selection should reflect the consequences of the decisions made following risk stratification. Third, we emphasize that there is usually no universally optimal threshold but rather that a plausible risk threshold depends on the clinical context. Consequently, we recommend to present results for multiple risk thresholds when developing or validating a prediction model. Conclusion Bearing in mind these three considerations can avoid inappropriate allocation (and non-allocation) of interventions. Using discriminating and well-calibrated models will generate better clinical outcomes if context-dependent thresholds are used.