Transcriptome network component analysis with limited microarray data

Transcriptome network component analysis with limited microarray data
复制标题

DOI:
10.1093/bioinformatics/btl279
复制
发表时间:
2006-08-01
期刊:
影响因子:
5.8
通讯作者:
Liao, James C.
Liao, James C.
中科院分区:
生物学3区
文献类型:
--
作者:
Galbraith, Simon J.;Tran, Linh M.;Liao, James C.

文献摘要

被引文献

相似文献

网络成分分析(Network Component Analysis,NCA)是一种从基因表达数据和转录因子(Transcription Factor,TF)-基因结合连接网络中推断TF活性和TF-基因调控强度的方法。以前,这种方法可以分析的监管机构的最大数量等于总样本量,因为在数据分解的可识别性限制。因此,源信号分量的总数限于实验的总数,而不是生物调节剂的总数。然而,具有比调节器的数量少的转录组数据点的网络是感兴趣的。因此,它是必要的,以发展一个理论基础,允许现实的源信号提取的基础上相对较少的数据点。另一方面,这种方法必然会增加数值挑战,导致多种解决方案。因此,这两个问题的解决方案是need.Results:我们已经改进了NCA的转录因子活性(TFA)的估计,根据观察,大多数基因的调控只有少数TF。这一观察导致推导出一个新的可识别性标准,该标准在数值迭代过程中进行测试,当TF的数量大于实验的数量时,该标准允许我们分解数据。为了表明我们的方法适用于真实的微阵列数据并具有生物学效用,我们使用来自ChIP芯片结合数据的TF-基因连接网络(96个TF)分析酿酒酵母细胞周期微阵列数据(73个实验)。我们比较了NCA分析的结果与从ChIP芯片回归方法获得的结果,我们表明,NCA和回归产生的TFA是定性相似的,但NCA TFA在统计检验中优于回归。我们还表明,NCA可以提取微妙的TFA信号,与已知的细胞周期TF功能和细胞周期相关联。总体而言,我们确定了31个TF在一个或多个实验中具有统计学周期性的TFA,其中75%是已知的细胞周期调节剂。此外,我们发现,在两个或多个实验中周期性的12个TFA对应于众所周知的细胞周期调节剂。我们还研究了TFA对连接网络选择的敏感性,我们使用不同的ChIP芯片p值截止值构建了两个网络。
Network component analysis (NCA) is a method to deduce transcription factor (TF) activities and TF-gene regulation control strengths from gene expression data and a TF-gene binding connectivity network. Previously, this method could analyze a maximum number of regulators equal to the total sample size because of the identifiability limit in data decomposition. As such, the total number of source signal components was limited to the total number of experiments rather than the total number of biological regulators. However, networks that have less transcriptome data points than the number of regulators are of interest. Thus it is imperative to develop a theoretical basis that allows realistic source signal extraction based on relatively few data points. On the other hand, such methods would inherently increase numerical challenges leading to multiple solutions. Therefore, solutions to both the problems are needed.Results: We have improved NCA for transcription factor activity (TFA) estimation, based on the observation that most genes are regulated by only a few TFs. This observation leads to the derivation of a new identifiability criterion which is tested during numerical iteration that allows us to decompose data when the number of TFs is greater than the number of experiments. To show that our method works with real microarray data and has biological utility, we analyze Saccharomyces cerevisiae cell cycle microarray data (73 experiments) using a TF-gene connectivity network (96 TFs) derived from ChIP-chip binding data. We compare the results of NCA analysis with the results obtained from ChIP-chip regression methods, and we show that NCA and regression produce TFAs that are qualitatively similar, but the NCA TFAs outperform regression in statistical tests. We also show that NCA can extract subtle TFA signals that correlate with known cell cycle TF function and cell cycle phase. Overall we determined that 31 TFs have statistically periodic TFAs in one or more experiments, 75% of which are known cell cycle regulators. In addition, we find that the 12 TFAs that are periodic in two or more experiments correspond to well-known cell cycle regulators. We also investigated TFA sensitivity to the choice of connectivity network we constructed two networks using different ChIP-chip p-value cut-offs.