Measuring the performance of prediction models to personalize treatment choice

Measuring the performance of prediction models to personalize treatment choice
复制标题

DOI:
10.1002/sim.9665
复制
发表时间:
2023-01-26
影响因子:
2
通讯作者:
White, Ian R.
White, Ian R.
中科院分区:
医学3区
文献类型:
--
作者:
Efthimiou, Orestis;Hoogland, Jeroen;White, Ian R.

文献摘要

被引文献

相似文献

当在随机试验中获得接受治疗或对照干预的个别患者的数据时,可以使用各种统计和机器学习方法来开发用于预测两种情况下的未来结果的模型,从而预测患者级别的治疗效果。这些预测随后可以指导个性化的治疗选择。虽然有几种方法可用于验证预测模型,但很少有人注意衡量个性化治疗效果预测的性能。在这篇文章中,我们提出了一系列可用于实现这一目标的措施。我们首先定义治疗效果和单一结果的模型准确性的两个维度:针对收益的区分和针对收益的校正。然后,我们将这两个维度合并为一个额外的概念,即决策准确性,它量化了模型识别从治疗中受益超过给定阈值的患者的能力。随后,我们提出了与这些维度相关的一系列绩效指标,并讨论了估计程序,重点是随机数据。我们的方法适用于连续或二元结果,适用于任何类型的预测模型,只要它使用基线协变量来预测治疗和控制下的结果。我们使用两个模拟数据集和一个抑郁症试验的真实数据集来说明所有方法。我们实现了R包Predival中的所有方法。结果表明,所提出的方法可用于评估和比较竞争模型在预测个体化治疗效果方面的表现。
When data are available from individual patients receiving either a treatment or a control intervention in a randomized trial, various statistical and machine learning methods can be used to develop models for predicting future outcomes under the two conditions, and thus to predict treatment effect at the patient level. These predictions can subsequently guide personalized treatment choices. Although several methods for validating prediction models are available, little attention has been given to measuring the performance of predictions of personalized treatment effect. In this article, we propose a range of measures that can be used to this end. We start by defining two dimensions of model accuracy for treatment effects, for a single outcome: discrimination for benefit and calibration for benefit. We then amalgamate these two dimensions into an additional concept, decision accuracy, which quantifies the model's ability to identify patients for whom the benefit from treatment exceeds a given threshold. Subsequently, we propose a series of performance measures related to these dimensions and discuss estimating procedures, focusing on randomized data. Our methods are applicable for continuous or binary outcomes, for any type of prediction model, as long as it uses baseline covariates to predict outcomes under treatment and control. We illustrate all methods using two simulated datasets and a real dataset from a trial in depression. We implement all methods in the R package predieval. Results suggest that the proposed measures can be useful in evaluating and comparing the performance of competing models in predicting individualized treatment effect.