Identification of Nonlinear State-Space Systems From Heterogeneous Datasets

Identification of Nonlinear State-Space Systems From Heterogeneous Datasets
复制标题

DOI:
10.1109/tcns.2017.2758966
复制
发表时间:
2018-06
影响因子:
4.2
通讯作者:
W. Pan;Ye Yuan;L. Ljung;J. Gonçalves;G. Stan
W. Pan;Ye Yuan;L. Ljung;J. Gonçalves;G. Stan
中科院分区:
计算机科学3区
文献类型:
--
作者:
W. Pan;Ye Yuan;L. Ljung;J. Gonçalves;G. Stan

文献摘要

被引文献

相似文献

本文提出了一种从异构数据集中识别非线性状态空间系统的新方法。该方法是在从实验数据中识别生化/基因网络(即识别反应动力学和动力学参数)的背景下描述的。同时集成各种数据集有可能为系统识别提供更好的性能。实验收集的数据通常因具体的实验设置和条件而异。通常,异质数据是通过实验获得的:1)从同一生物系统中重复测量或2)应用不同的实验条件,如生物诱导、温度、基因敲除、基因过表达等的变化/扰动。我们在这里使用贝叶斯学习框架来制定识别问题,该框架利用“稀疏组”先验来允许最稀疏模型的推理,该模型可以解释整个观察到的异构数据集。为了能够扩展到大量的特征,使用凸凹过程将得到的非凸优化问题松弛为一个重新加权的Group Lasso问题。作为我们方法有效性的一个例子,我们用它来识别一个遗传振荡器(广义八种再压器)。通过这个例子,我们表明,当实验次数增加时,即使单个时间序列数据很短,我们的算法也优于Group Lasso。我们还通过改变过程噪声和测量噪声的强度来评估我们的算法对噪声的鲁棒性。
This paper proposes a new method to identify nonlinear state-space systems from heterogeneous datasets. The method is described in the context of identifying biochemical/gene networks (i.e., identifying both reaction dynamics and kinetic parameters) from experimental data. Simultaneous integration of various datasets has the potential to yield better performance for system identification. Data collected experimentally typically vary depending on the specific experimental setup and conditions. Typically, heterogeneous data are obtained experimentally through 1) replicate measurements from the same biological system or 2) application of different experimental conditions such as changes/perturbations in biological inductions, temperature, gene knock-out, gene over-expression, etc. We formulate here the identification problem using a Bayesian learning framework that makes use of “sparse group” priors to allow inference of the sparsest model that can explain the whole set of observed heterogeneous data. To enable scale up to large number of features, the resulting nonconvex optimization problem is relaxed to a reweighted Group Lasso problem using a convex–concave procedure. As an illustrative example of the effectiveness of our method, we use it to identify a genetic oscillator (generalized eight species repressilator). Through this example we show that our algorithm outperforms Group Lasso when the number of experiments is increased, even when each single time-series dataset is short. We additionally assess the robustness of our algorithm against noise by varying the intensity of process noise and measurement noise.