Identifying context-specific transcription factor targets from prior knowledge and gene expression data.

Identifying context-specific transcription factor targets from prior knowledge and gene expression data.
复制标题

DOI:
10.1109/tnb.2013.2263390
复制
发表时间:
2013-09
影响因子:
3.9
通讯作者:
Ochs MF
Ochs MF
中科院分区:
生物学3区
文献类型:
--
作者:
Fertig EJ;Favorov AV;Ochs MF

文献摘要

被引文献

相似文献

目前,许多方法、测定和数据库提供了转录因子(TF)的候选靶标。然而,信托基金很少普遍地调节其目标。TF激活的背景可以改变靶的转录应答。哺乳动物基因典型的直接多重调控使从基因表达数据直接推断TF靶点复杂化。我们提出了一种新的统计推断上下文特定的TF调控的基础上的CoGAPS算法,推断重叠的基因表达模式,导致共调节。模拟数据的数值实验表明,这种统计正确推断的目标是共同的多个TF,除了在从TF的信号是可以忽略不计的情况下,相对于噪声水平和信号从其他TF。该统计对模拟基因集中的中等水平的错误是稳健的,识别出比假阴性更少的假阳性。值得注意的是,监管统计数据将与胃肠道间质瘤(GIST)中细胞信号传导相关的TF靶点数量细化为与先前研究中鉴定的TF磷酸化模式一致的基因。作为制定,拟议的监管统计数据具有广泛的适用性,推断集成数据集的成员资格。该统计量可以自然地扩展到考虑集合成员资格的先验概率或添加候选基因靶。
Numerous methodologies, assays, and databases presently provide candidate targets of transcription factors (TFs). However, TFs rarely regulate their targets universally. The context of activation of a TF can change the transcriptional response of targets. Direct multiple regulation typical to mammalian genes complicates direct inference of TF targets from gene expression data. We present a novel statistic that infers context-specific TF regulation based upon the CoGAPS algorithm, which infers overlapping gene expression patterns resulting from coregulation. Numerical experiments with simulated data showed that this statistic correctly inferred targets that are common to multiple TFs, except in cases where the signal from a TF is negligible relative to noise level and signal from other TFs. The statistic is robust to moderate levels of error in the simulated gene sets, identifying fewer false positives than false negatives. Significantly, the regulatory statistic refines the number of TF targets relevant to cell signaling in gastrointestinal stromal tumors (GIST) to genes consistent with the phosphorylation patterns of TFs identified in previous studies. As formulated, the proposed regulatory statistic has wide applicability to inferring set membership in integrated datasets. This statistic could be naturally extended to account for prior probabilities of set membership or to add candidate gene targets.