Utilizing the Heterogeneity of Clinical Data for Model Refinement and Rule Discovery Through the Application of Genetic Algorithms to Calibrate a High-Dimensional Agent-Based Model of Systemic Inflammation.

Utilizing the Heterogeneity of Clinical Data for Model Refinement and Rule Discovery Through the Application of Genetic Algorithms to Calibrate a High-Dimensional Agent-Based Model of Systemic Inflammation.
复制标题

DOI:
10.3389/fphys.2021.662845
复制
发表时间:
2021
影响因子:
4
通讯作者:
An G
An G
中科院分区:
医学2区
文献类型:
--
作者:
Cockrell C;An G

文献摘要

参考文献

被引文献

相似文献

生物异质性的解释是生物医学研究中最大的挑战之一。动态计算和数学模型可用于增强对生物系统的研究和理解,但是用于校准和验证的传统方法通常不考虑生物数据的异质性,这可能导致这些模型的过拟合和脆性。在本文中,我们提出了一种机器学习方法,该方法利用遗传算法(GA)来校准和改进急性全身性炎症的基于代理的模型(ABM),重点是考虑临床数据集中的异质性,从而避免过拟合并提高基础模拟模型的鲁棒性和潜在的可推广性。方法:基于Agent的建模方法是多尺度机理建模中常用的建模方法。然而,同样的属性,使ABM非常适合代表生物系统也提出了重大的挑战,其建设和校准方面,特别是在选择潜在的机械规则和大量的相关自由参数。我们已经提出,机器学习方法(如遗传算法)可以用来更有效地处理规则选择和参数空间表征;目前的工作适用于遗传算法的挑战,校准一个复杂的ABM到一个特定的数据集,同时保持生物异质性反映在范围和方差的数据。该项目使用GA来增强先前验证的急性全身性炎症的ABM的规则集,先天免疫应答ABM(IIRABM)到来自烧伤患者群体的全身性细胞因子水平的临床时间序列数据。GA的基因组是从IIRABM的模型规则矩阵(MRM)生成的向量,该矩阵不仅表示与IIRABM的细胞因子相互作用规则相关的常数/参数,还表示规则本身的存在。通过结合临床数据的样本值范围(“误差条”)的适应度函数来捕获异质性。结果如下:GA使能的参数空间探索导致一组推定的MRM规则和相关参数化,其紧密匹配用于设计适应度函数的细胞因子时程数据。随着模型参数化朝着适应度函数最小值发展,MRM中非零元素的数量显著增加,从稀疏矩阵过渡到密集矩阵。这导致模型结构更接近(在表面水平上)由标准差异基因表达实验研究产生的数据结构。结论:我们提出了一种支持HPC的机器学习/进化计算方法,用于将复杂的ABM校准到复杂的临床数据,同时保留生物异质性。机器学习、HPC和多尺度机制建模的集成为更有效地表示临床人群及其数据的异质性提供了一条途径。
Introduction: Accounting for biological heterogeneity represents one of the greatest challenges in biomedical research. Dynamic computational and mathematical models can be used to enhance the study and understanding of biological systems, but traditional methods for calibration and validation commonly do not account for the heterogeneity of biological data, which may result in overfitting and brittleness of these models. Herein we propose a machine learning approach that utilizes genetic algorithms (GAs) to calibrate and refine an agent-based model (ABM) of acute systemic inflammation, with a focus on accounting for the heterogeneity seen in a clinical data set, thereby avoiding overfitting and increasing the robustness and potential generalizability of the underlying simulation model. Methods: Agent-based modeling is a frequently used modeling method for multi-scale mechanistic modeling. However, the same properties that make ABMs well suited to representing biological systems also present significant challenges with respect to their construction and calibration, particularly with respect to the selection of potential mechanistic rules and the large number of associated free parameters. We have proposed that machine learning approaches (such as GAs) can be used to more effectively and efficiently deal with rule selection and parameter space characterization; the current work applies GAs to the challenge of calibrating a complex ABM to a specific data set, while preserving biological heterogeneity reflected in the range and variance of the data. This project uses a GA to augment the rule-set for a previously validated ABM of acute systemic inflammation, the Innate Immune Response ABM (IIRABM) to clinical time series data of systemic cytokine levels from a population of burn patients. The genome for the GA is a vector generated from the IIRABM’s Model Rule Matrix (MRM), which is a matrix representation of not only the constants/parameters associated with the IIRABM’s cytokine interaction rules, but also the existence of rules themselves. Capturing heterogeneity is accomplished by a fitness function that incorporates the sample value range (“error bars”) of the clinical data. Results: The GA-enabled parameter space exploration resulted in a set of putative MRM rules and associated parameterizations which closely match the cytokine time course data used to design the fitness function. The number of non-zero elements in the MRM increases significantly as the model parameterizations evolve toward a fitness function minimum, transitioning from a sparse to a dense matrix. This results in a model structure that more closely resembles (at a superficial level) the structure of data generated by a standard differential gene expression experimental study. Conclusion: We present an HPC-enabled machine learning/evolutionary computing approach to calibrate a complex ABM to complex clinical data while preserving biological heterogeneity. The integration of machine learning, HPC, and multi-scale mechanistic modeling provides a pathway forward to more effectively representing the heterogeneity of clinical populations and their data.
DOI: 10.1371/journal.pone.0122192
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Cockrell RC;Christley S;Chang E;An G
通讯作者: An G
DOI: 10.1002/ddr.20415
发表时间: 2011-03-01
影响因子: 3.8
作者:
An, Gary;Bartels, John;Vodovotz, Yoram
通讯作者: Vodovotz, Yoram
DOI: 10.1016/j.burns.2018.09.001
发表时间: 2019-03-01
期刊: BURNS
影响因子: 2.7
作者:
Bergquist, Maria;Hastbacka, Johanna;Lipcsey, Miklos
通讯作者: Lipcsey, Miklos
DOI: 10.1177/2472555216682725
发表时间: 2017-03
期刊: SLAS discovery : advancing life sciences R & D
影响因子: --
作者:
Gough A;Stern AM;Maier J;Lezon T;Shun TY;Chennubhotla C;Schurdak ME;Haney SA;Taylor DL
通讯作者: Taylor DL
DOI: 10.1016/j.ress.2005.11.014
发表时间: 2006-10-01
影响因子: 8.1
作者:
Saltelli, Andrea;Ratto, Marco;Campolongo, Francesca
通讯作者: Campolongo, Francesca