Machine learning methods for leveraging baseline covariate information to improve the efficiency of clinical trials

Machine learning methods for leveraging baseline covariate information to improve the efficiency of clinical trials
复制标题

利用基线协变量信息提高临床试验效率的机器学习方法

DOI:
10.1002/sim.8054
复制
发表时间:
2019
影响因子:
2
通讯作者:
Ma, Shujie
Ma, Shujie
中科院分区:
医学3区
文献类型:
--
作者:
Zhang, Zhiwei;Ma, Shujie

文献摘要

参考文献

被引文献

相似文献

临床试验被广泛认为是治疗评估的黄金标准,但在时间和金钱方面可能非常昂贵。通过整合与临床结果相关的基线协变量的信息可以提高临床试验的效率。这可以通过使用涉及协变量函数的增强项修改未调整的治疗效果估计器来完成。最佳增强在理论上已得到很好的描述,但必须在实践中进行估计。在本文中,我们研究使用机器学习方法来估计最佳增强。我们考虑并比较基于估计回归函数的间接方法和旨在直接最小化治疗效果估计量的渐近方差的直接方法。理论考虑和模拟结果表明直接方法通常优于间接方法。可以使用可以最小化预测误差加权平方和的任何现有预测算法来实现直接方法。这样的预测算法有很多,并且可以利用超级学习原理,在直接方法下将多种算法组合成一个超级学习器。由此产生的直接超级学习器具有理想的预言机属性,易于实现,并且在现实环境中表现良好。所提出的方法通过中风试验的真实数据进行说明。
Clinical trials are widely considered the gold standard for treatment evaluation, and they can be highly expensive in terms of time and money. The efficiency of clinical trials can be improved by incorporating information from baseline covariates that are related to clinical outcomes. This can be done by modifying an unadjusted treatment effect estimator with an augmentation term that involves a function of covariates. The optimal augmentation is well characterized in theory but must be estimated in practice. In this article, we investigate the use of machine learning methods to estimate the optimal augmentation. We consider and compare an indirect approach based on an estimated regression function and a direct approach that aims directly to minimize the asymptotic variance of the treatment effect estimator. Theoretical considerations and simulation results indicate that the direct approach is generally preferable over the indirect approach. The direct approach can be implemented using any existing prediction algorithm that can minimize a weighted sum of squared prediction errors. Many such prediction algorithms are available, and the super learning principle can be used to combine multiple algorithms into a super learner under the direct approach. The resulting direct super learner has a desirable oracle property, is easy to implement, and performs well in realistic settings. The proposed methodology is illustrated with real data from a stroke trial.
使用广义线性模型来利用基线变量,对随机试验中的治疗效果进行简单、有效的估计。
DOI: 10.2202/1557-4679.1138
发表时间: 2010
期刊: The international journal of biostatistics
影响因子: --
作者:
Rosenblum,Michael;vanderLaan,MarkJ
通讯作者: vanderLaan,MarkJ
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者:
Clare Wilson
通讯作者: Clare Wilson
DOI: 10.1093/biostatistics/kxr050
发表时间: 2012-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Tian, Lu;Cai, Tianxi;Wei, Lee-Jen
通讯作者: Wei, Lee-Jen