Optimizing the synthesis of clinical trial data using sequential trees.

Optimizing the synthesis of clinical trial data using sequential trees.
复制标题

DOI:
10.1093/jamia/ocaa249
复制
发表时间:
2021-01-15
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Zheng C
Zheng C
中科院分区:
其他
文献类型:
--
作者:
Emam KE;Mosquera L;Zheng C

文献摘要

参考文献

被引文献

相似文献

随着对共享临床试验数据的需求不断增长,需要可扩展的方法来实现对高效用数据的隐私保护访问。数据合成就是这样一种方法。序列树通常用于合成健康数据。假设所生成的数据的效用取决于变量顺序。迄今为止,尚未评估变量顺序对合成临床试验数据的影响。通过模拟,我们的目标是评估的可变性,在效用的合成临床试验数据的变量顺序随机洗牌,并实现优化算法,以找到一个良好的顺序,如果变异性太高。在模拟中评估了六个肿瘤学临床试验数据集。三个效用度量计算比较真实的和合成的数据:单变量相似性,相似性在多变量预测准确性,和可重复性度量。利用粒子群算法优化变量排序,并与课程学习算法进行比较。随着临床试验数据集中变量数量的增加,数据效用的变异性随顺序显著增加。粒子群与可重复性铰链损失确保了足够的效用在所有6个数据集。选择铰链阈值是为了避免过度拟合,过度拟合会造成隐私问题。从实用性来看,这比课程学习要优越上级。本研究提出的优化方法提供了一种可靠的方法来合成高效用的临床试验数据集。
With the growing demand for sharing clinical trial data, scalable methods to enable privacy protective access to high-utility data are needed. Data synthesis is one such method. Sequential trees are commonly used to synthesize health data. It is hypothesized that the utility of the generated data is dependent on the variable order. No assessments of the impact of variable order on synthesized clinical trial data have been performed thus far. Through simulation, we aim to evaluate the variability in the utility of synthetic clinical trial data as variable order is randomly shuffled and implement an optimization algorithm to find a good order if variability is too high. Six oncology clinical trial datasets were evaluated in a simulation. Three utility metrics were computed comparing real and synthetic data: univariate similarity, similarity in multivariate prediction accuracy, and a distinguishability metric. Particle swarm was implemented to optimize variable order, and was compared with a curriculum learning approach to ordering variables. As the number of variables in a clinical trial dataset increases, there is a pattern of a marked increase in variability of data utility with order. Particle swarm with a distinguishability hinge loss ensured adequate utility across all 6 datasets. The hinge threshold was selected to avoid overfitting which can create a privacy problem. This was superior to curriculum learning in terms of utility. The optimization approach presented in this study gives a reliable way to synthesize high-utility clinical trial datasets.
DOI: 10.1038/srep01376
发表时间: 2013
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
de Montjoye, Yves-Alexandre;Hidalgo, Cesar A.;Verleysen, Michel;Blondel, Vincent D.
通讯作者: Blondel, Vincent D.
DOI: 10.1037/pspp0000208
发表时间: 2021-08-01
影响因子: 7.6
作者:
Arslan, Ruben C.;Schilling, Katharina M.;Penke, Lars
通讯作者: Penke, Lars
DOI: 10.1001/jama.2014.9646
发表时间: 2014-09-10
影响因子: 120.7
作者:
Ebrahim, Shanil;Sohani, Zahra N.;Ioannidis, John P. A.
通讯作者: Ioannidis, John P. A.
DOI: 10.1214/07-sts242
发表时间: 2007-11-01
影响因子: 5.7
作者:
Buehlmann, Peter;Hothorn, Torsten
通讯作者: Hothorn, Torsten
DOI: 10.2903/j.efsa.2019.5926
发表时间: 2019-12-01
期刊: EFSA JOURNAL
影响因子: 3.3
作者:
通讯作者: --