SynTwin: A graph-based approach for predicting clinical outcomes using digital twins derived from synthetic patients

SynTwin: A graph-based approach for predicting clinical outcomes using digital twins derived from synthetic patients
复制标题

DOI:
10.1142/9789811286421_0008
复制
发表时间:
2023-12
影响因子:
--
通讯作者:
Jason H Moore;Xi Li;Jui-Hsuan Chang;Nicholas P. Tatonetti;Dan Theodorescu;Yong Chen;F. Asselbergs;Mythreye Venkatesan;Zhiping Wang
Jason H Moore;Xi Li;Jui-Hsuan Chang;Nicholas P. Tatonetti;Dan Theodorescu;Yong Chen;F. Asselbergs;Mythreye Venkatesan;Zhiping Wang
中科院分区:
--
文献类型:
--
作者:
Jason H Moore;Xi Li;Jui-Hsuan Chang;Nicholas P. Tatonetti;Dan Theodorescu;Yong Chen;F. Asselbergs;Mythreye Venkatesan;Zhiping Wang

文献摘要

相似文献

数字孪生的概念来自工程、工业和制造领域,用于创建虚拟对象或机器,从而为真实的对象的设计和开发提供信息。这一想法对精准医学很有吸引力,因为患者的数字双胞胎可以帮助为医疗决策提供信息。我们开发了一种方法来生成和使用数字双胞胎进行临床结果预测。我们引入了一种新的方法,将合成数据和网络科学相结合,为精准医疗创建数字孪生(即SynTwin)。首先,我们的方法首先根据所有对象的可用特征估计它们之间的距离。其次,距离被用来构建一个网络与主题作为节点和边缘定义的距离小于渗透阈值。第三,主体的社区或集团被定义。第四,使用合成数据生成算法生成大量合成患者,该合成数据生成算法对数据的相关结构进行建模以生成新患者。第五,数字双胞胎是从在网络中定义受试者社区的给定距离内的合成患者群体中选择的。最后,我们使用真实的受试者、数字双胞胎或社区内外的受试者,比较和对比基于社区的临床终点预测。这种方法的关键是使用患者相似性定义的数字双胞胎,其表示具有与由网络距离和社区结构定义的附近真实的患者相似的模式的假设未观察到的患者。我们应用我们的SynTwin方法预测来自美国国家癌症研究所(USA)的监测、流行病学和最终结果(SEER)计划的基于人群的癌症登记(n= 87,674)中的死亡率。我们的研究结果表明,在这项研究中,数字双胞胎的最近网络邻居预测死亡率(AUROC=0.864,95%CI =0.857-0.872)比仅使用真实的数据(AUROC=0.791,95%CI =0.781-0.800)有显著改善。这些结果表明,使用合成患者的基于网络的数字孪生策略可能会为精准医疗工作增加价值。
The concept of a digital twin came from the engineering, industrial, and manufacturing domains to create virtual objects or machines that could inform the design and development of real objects. This idea is appealing for precision medicine where digital twins of patients could help inform healthcare decisions. We have developed a methodology for generating and using digital twins for clinical outcome prediction. We introduce a new approach that combines synthetic data and network science to create digital twins (i.e. SynTwin) for precision medicine. First, our approach starts by estimating the distance between all subjects based on their available features. Second, the distances are used to construct a network with subjects as nodes and edges defining distance less than the percolation threshold. Third, communities or cliques of subjects are defined. Fourth, a large population of synthetic patients are generated using a synthetic data generation algorithm that models the correlation structure of the data to generate new patients. Fifth, digital twins are selected from the synthetic patient population that are within a given distance defining a subject community in the network. Finally, we compare and contrast community-based prediction of clinical endpoints using real subjects, digital twins, or both within and outside of the community. Key to this approach are the digital twins defined using patient similarity that represent hypothetical unobserved patients with patterns similar to nearby real patients as defined by network distance and community structure. We apply our SynTwin approach to predicting mortality in a population-based cancer registry (n=87,674) from the Surveillance, Epidemiology, and End Results (SEER) program from the National Cancer Institute (USA). Our results demonstrate that nearest network neighbor prediction of mortality in this study is significantly improved with digital twins (AUROC=0.864, 95% CI=0.857–0.872) over just using real data alone (AUROC=0.791, 95% CI=0.781–0.800). These results suggest a network-based digital twin strategy using synthetic patients may add value to precision medicine efforts.