A model-based optimization framework for the inference of regulatory interactions using time-course DNA microarray expression data.

A model-based optimization framework for the inference of regulatory interactions using time-course DNA microarray expression data.
复制标题

DOI:
10.1186/1471-2105-8-228
复制
发表时间:
2007-06-29
期刊:
影响因子:
3
通讯作者:
Papoutsakis, Eleftherios T
Papoutsakis, Eleftherios T
中科院分区:
生物学4区
文献类型:
--
作者:
Thomas, Reuben;Paredes, Carlos J;Mehrotra, Sanjay;Hatzimanikatis, Vassily;Papoutsakis, Eleftherios T

文献摘要

相似文献

蛋白质是转录的主要调节因子,尽管来自DNA微阵列等系统的mRNA表达数据本身已被广泛使用。此外,遗传系统中的调控过程本质上是非线性的,大多数研究采用mRNA表达的时程分析。这些考虑因素应考虑到在遗传网络中的调节相互作用的推断方法的发展。我们使用基于S系统的模型进行转录和翻译过程。我们提出了一种基于优化的调控网络推理方法,使用来自DNA微阵列分析的时变数据。目前,这似乎是唯一的基于模型的方法,可用于分析时间过程的“相对”表达(表达率)。我们进行了分析的动态行为的系统时,可用的实验样本的数量是不同的,当有不同程度的噪音在数据中,当有基因不被实验者考虑。我们的研究表明,影响一种方法正确推断相互作用的能力的主要因素是一些或所有基因的时间曲线的相似性。简档彼此越不相似,就越容易推断出相互作用。我们提出了一个启发式的方法来解决网络,并表明它显示出合理的性能上的合成网络。最后,我们验证了我们的方法使用真实的实验数据的一个选定的子集的基因参与炭疽芽孢杆菌的孢子形成级联。我们表明,该方法捕获了所选基因之间的大多数重要的已知相互作用。基因间调控相互作用的任何推理方法的性能取决于数据中的噪声、影响网络基因的未知基因的存在以及部分或全部基因的时间曲线的相似性。虽然受到这些问题,本文中提出的推理方法将是有用的,因为它能够推断重要的相互作用,事实上,它可以与时程DNA微阵列数据,因为它是基于一个非线性模型的过程,明确占蛋白质的调节作用。
Proteins are the primary regulatory agents of transcription even though mRNA expression data alone, from systems like DNA microarrays, are widely used. In addition, the regulation process in genetic systems is inherently non-linear in nature, and most studies employ a time-course analysis of mRNA expression. These considerations should be taken into account in the development of methods for the inference of regulatory interactions in genetic networks. We use an S-system based model for the transcription and translation process. We propose an optimization-based regulatory network inference approach that uses time-varying data from DNA microarray analysis. Currently, this seems to be the only model-based method that can be used for the analysis of time-course "relative" expressions (expression ratios). We perform an analysis of the dynamic behavior of the system when the number of experimental samples available is varied, when there are different levels of noise in the data and when there are genes that are not considered by the experimenter. Our studies show that the principal factor affecting the ability of a method to infer interactions correctly is the similarity in the time profiles of some or all the genes. The less similar the profiles are to each other the easier it is to infer the interactions. We propose a heuristic method for resolving networks and show that it displays reasonable performance on a synthetic network. Finally, we validate our approach using real experimental data for a chosen subset of genes involved in the sporulation cascade of Bacillus anthracis. We show that the method captures most of the important known interactions between the chosen genes. The performance of any inference method for regulatory interactions between genes depends on the noise in the data, the existence of unknown genes affecting the network genes, and the similarity in the time profiles of some or all genes. Though subject to these issues, the inference method proposed in this paper would be useful because of its ability to infer important interactions, the fact that it can be used with time-course DNA microarray data and because it is based on a non-linear model of the process that explicitly accounts for the regulatory role of proteins.