A machine learning approach for identifying novel cell type-specific transcriptional regulators of myogenesis.

A machine learning approach for identifying novel cell type-specific transcriptional regulators of myogenesis.
复制标题

DOI:
10.1371/journal.pgen.1002531
复制
发表时间:
2012
期刊:
影响因子:
4.5
通讯作者:
Michelson AM
Michelson AM
中科院分区:
生物学2区
文献类型:
--
作者:
Busser BW;Taher L;Kim Y;Tansey T;Bloom MJ;Ovcharenko I;Michelson AM

文献摘要

参考文献

被引文献

相似文献

转录增强子整合了多种转录因子(TF)的贡献,以协调发育过程中发生的无数时空基因表达程序。具有相似活性的增强子的分子理解需要鉴定它们的独特和共享的序列特征。为了解决这个问题,我们结合系统发育分析与基于DNA的增强子序列分类器,分析TF结合位点(TFBS)管理的共表达基因集的转录。我们首先组装了少量的增强子,这些增强子在果蝇肌肉创始细胞(FC)和其他中胚层细胞类型中具有活性。使用系统发育分析,我们增加了增强子的数量,将正交,但不同的序列从其他果蝇物种。功能分析显示,趋异增强子直系同源物的活性模式与它们的D。黑腹动物的对应物,尽管已知的TFBS存在广泛的进化洗牌。然后,我们使用该增强子集合构建并训练分类器,并基于已知和推定的TFBS的存在或不存在来鉴定另外的相关增强子。预测的FC增强子在已知FC基因附近过度表达;并且发现通过分类器学习的许多TFBS对增强子活性至关重要,包括POU同源结构域、Myb、Ets、Forkhead和T-box基序。经验测试还表明,由org-1编码的T盒TF是以前未表征的肌细胞身份调节剂。最后,我们发现了广泛的多样性,在已知的FC增强子的TFBS的组成,这表明基序组合在这种增强子所表现出的细胞特异性中起着至关重要的作用。总之,机器学习结合进化序列分析可用于识别新型TFBS,并有助于鉴定协调细胞类型特异性发育基因表达模式的同源TF。多细胞生物的发育需要形成多种细胞类型。每个细胞都有一个独特的遗传程序,该程序由称为增强子的调控序列编排,该序列包含多个结合不同转录因子的短DNA序列。了解发育调控网络需要了解功能相关增强子的序列特征。我们开发了一种综合的进化和计算方法来破译增强子调控代码,并应用这种方法来发现控制果蝇肌肉发育的转录网络的新组件。我们的方法包括组装已知的肌肉增强子,用进化上保守的序列扩展这一组,根据它们的共享序列特征对这些增强子进行计算分类,并扫描整个果蝇基因组以预测其他相关的增强子。使用这种方法,我们创建了5,500个推定的肌肉增强子的图谱,确定了它们结合的候选转录因子,观察到映射的增强子和肌肉基因表达之间的强相关性,并发现了验证的肌肉增强子中转录因子结合位点组合之间的广泛异质性,这一特征可能有助于这些调控元件的个体细胞特异性。我们的策略可以很容易地推广到研究其他生物和发育背景下的转录网络。
Transcriptional enhancers integrate the contributions of multiple classes of transcription factors (TFs) to orchestrate the myriad spatio-temporal gene expression programs that occur during development. A molecular understanding of enhancers with similar activities requires the identification of both their unique and their shared sequence features. To address this problem, we combined phylogenetic profiling with a DNA–based enhancer sequence classifier that analyzes the TF binding sites (TFBSs) governing the transcription of a co-expressed gene set. We first assembled a small number of enhancers that are active in Drosophila melanogaster muscle founder cells (FCs) and other mesodermal cell types. Using phylogenetic profiling, we increased the number of enhancers by incorporating orthologous but divergent sequences from other Drosophila species. Functional assays revealed that the diverged enhancer orthologs were active in largely similar patterns as their D. melanogaster counterparts, although there was extensive evolutionary shuffling of known TFBSs. We then built and trained a classifier using this enhancer set and identified additional related enhancers based on the presence or absence of known and putative TFBSs. Predicted FC enhancers were over-represented in proximity to known FC genes; and many of the TFBSs learned by the classifier were found to be critical for enhancer activity, including POU homeodomain, Myb, Ets, Forkhead, and T-box motifs. Empirical testing also revealed that the T-box TF encoded by org-1 is a previously uncharacterized regulator of muscle cell identity. Finally, we found extensive diversity in the composition of TFBSs within known FC enhancers, suggesting that motif combinatorics plays an essential role in the cellular specificity exhibited by such enhancers. In summary, machine learning combined with evolutionary sequence analysis is useful for recognizing novel TFBSs and for facilitating the identification of cognate TFs that coordinate cell type–specific developmental gene expression patterns. The development of multicellular organisms requires the formation of a diversity of cell types. Each cell has a unique genetic program that is orchestrated by regulatory sequences called enhancers, comprising multiple short DNA sequences that bind distinct transcription factors. Understanding developmental regulatory networks requires knowledge of the sequence features of functionally related enhancers. We developed an integrated evolutionary and computational approach for deciphering enhancer regulatory codes and applied this method to discover new components of the transcriptional network controlling muscle development in the fruit fly, Drosophila melanogaster. Our method involves assembling known muscle enhancers, expanding this set with evolutionarily conserved sequences, computationally classifying these enhancers based on their shared sequence features, and scanning the entire Drosophila genome to predict additional related enhancers. Using this approach, we created a map of 5,500 putative muscle enhancers, identified candidate transcription factors to which they bind, observed a strong correlation between mapped enhancers and muscle gene expression, and uncovered extensive heterogeneity among combinations of transcription factor binding sites in validated muscle enhancers, a feature that may contribute to the individual cellular specificities of these regulatory elements. Our strategy can readily be generalized to study transcriptional networks in other organisms and developmental contexts.
调节性DNA序列的某些统计特性及其在预测果蝇基因组中的调节区域的使用:蓬松的尾巴测试。
DOI: 10.1186/1471-2105-6-109
发表时间: 2005-04-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Abnizova, I;te Boekhorst, R;Walter, K;Gilks, WR
通讯作者: Gilks, WR
DOI: 10.1073/pnas.231608898
发表时间: 2002-01-22
影响因子: 11.1
作者:
Berman, BP;Nibu, Y;Eisen, MB
通讯作者: Eisen, MB
DOI: 10.1006/dbio.2002.0606
发表时间: 2002-04-15
影响因子: 2.7
作者:
Carmena, A;Buff, E;Michelson, AM
通讯作者: Michelson, AM
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1016/j.cell.2008.05.024
发表时间: 2008-06-27
期刊: CELL
影响因子: 64.5
作者:
Berger, Michael F.;Badis, Gwenael;Hughes, Timothy R.
通讯作者: Hughes, Timothy R.