Machine learning enabled subgroup analysis with real-world data to inform clinical trial eligibility criteria design.

Machine learning enabled subgroup analysis with real-world data to inform clinical trial eligibility criteria design.
复制标题

DOI:
10.1038/s41598-023-27856-1
复制
发表时间:
2023-01-12
期刊:
影响因子:
4.6
通讯作者:
--
中科院分区:
综合性期刊3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

过度限制的临床试验合格性标准可能会限制试验结果对目标现实世界患者人群的普遍性。我们开发了一种新的机器学习方法,使用大量真实世界数据(RWD)来更好地为临床试验合格性标准设计提供信息。我们从电子健康记录(EHR)中提取了患者的临床事件,其中包括人口统计学,诊断和药物,并假设这些临床事件在个体EHR中的某些组成可以确定患者的亚表型同质集群,其中每个亚组中的患者具有相似的临床特征。我们引入了一个结果导向的概率模型来识别这些亚表型,使得同一亚组中的患者不仅具有相似的临床特征,而且遇到严重不良事件(SAE)的风险水平也相似。我们在之前使用OneFlorida+临床研究联盟的EHR进行的两项临床试验中评估了我们的算法。我们的模型可以以透明和可解释的方式清楚地识别更可能患有或不患有SAE的患者亚组作为亚表型。我们的方法确定了一组临床主题,并根据它们导出新的患者表示。每个临床主题表示从患者EHR学习的特定临床事件组成模式。在两项试验中进行测试,患者亚组(#SAE=0)和患者亚组(#SAE>0)可以通过使用推断主题的k均值聚类很好地分离。推断出的主题被表征为可能与患者亚组(#SAE>0)一致,揭示了临床特征的有意义的组合,并可以为完善临床试验的排除标准提供数据驱动的建议。所提出的监督主题建模方法可以从有或无SAE的亚表型推断临床主题。可以进一步推导出描述发生SAE的患者亚组的潜在规则,以告知临床试验合格性标准的设计。
Overly restrictive eligibility criteria for clinical trials may limit the generalizability of the trial results to their target real-world patient populations. We developed a novel machine learning approach using large collections of real-world data (RWD) to better inform clinical trial eligibility criteria design. We extracted patients’ clinical events from electronic health records (EHRs), which include demographics, diagnoses, and drugs, and assumed certain compositions of these clinical events within an individual’s EHRs can determine the subphenotypes—homogeneous clusters of patients, where patients within each subgroup share similar clinical characteristics. We introduced an outcome-guided probabilistic model to identify those subphenotypes, such that the patients within the same subgroup not only share similar clinical characteristics but also at similar risk levels of encountering severe adverse events (SAEs). We evaluated our algorithm on two previously conducted clinical trials with EHRs from the OneFlorida+ Clinical Research Consortium. Our model can clearly identify the patient subgroups who are more likely to suffer or not suffer from SAEs as subphenotypes in a transparent and interpretable way. Our approach identified a set of clinical topics and derived novel patient representations based on them. Each clinical topic represents a certain clinical event composition pattern learned from the patient EHRs. Tested on both trials, patient subgroup (#SAE=0) and patient subgroup (#SAE>0) can be well-separated by k-means clustering using the inferred topics. The inferred topics characterized as likely to align with the patient subgroup (#SAE>0) revealed meaningful combinations of clinical features and can provide data-driven recommendations for refining the exclusion criteria of clinical trials. The proposed supervised topic modeling approach can infer the clinical topics from the subphenotypes with or without SAEs. The potential rules for describing the patient subgroups with SAEs can be further derived to inform the design of clinical trial eligibility criteria.
DOI: 10.1200/jco.2017.73.7916
发表时间: 2017-11-20
影响因子: 45.3
作者:
Kim, Edward S.;Bruinooge, Suanna S.;Schilsky, Richard L.
通讯作者: Schilsky, Richard L.
DOI: 10.1186/s13195-016-0201-2
发表时间: 2016-08-12
期刊: Alzheimer's research & therapy
影响因子: --
作者:
Banzi R;Camaioni P;Tettamanti M;Bertele' V;Lucca U
通讯作者: Lucca U
DOI: 10.1162/jmlr.2003.3.4-5.993
发表时间: 2003-05-15
影响因子: 6
作者:
Blei, DM;Ng, AY;Jordan, MI
通讯作者: Jordan, MI
DOI: 10.1214/aoms/1177730491
发表时间: 1947-01-01
影响因子: --
作者:
MANN, HB;WHITNEY, DR
通讯作者: WHITNEY, DR
DOI: 10.1207/s15324796abm2902s_5
发表时间: 2005-01-01
影响因子: 3.8
作者:
Ory, M;Resnick, B;Bazzarre, T
通讯作者: Bazzarre, T