Sequence features of DNA binding sites reveal structural class of associated transcription factor

Sequence features of DNA binding sites reveal structural class of associated transcription factor
复制标题

DOI:
10.1093/bioinformatics/bti731
复制
发表时间:
2006-01-15
期刊:
影响因子:
5.8
通讯作者:
Hartemink, AJ
Hartemink, AJ
中科院分区:
生物学3区
文献类型:
--
作者:
Narlikar, L;Hartemink, AJ

文献摘要

被引文献

相似文献

动机:分子生物学的一个关键目标是了解细胞调节其基因转录的机制。这种转录调控的一个重要方面是转录因子(tf)与它们在DNA上的特定顺式调控对应物的结合。tf根据其DNA结合结构域(如锌指、亮氨酸拉链、同源结构域)的结构识别和结合其DNA对应物。这些域的结构可以用作将tf分组为类的基础。尽管DNA结合域的结构在不同的tf中差异很大,但特定类别的tf以相似的方式与DNA结合,这表明在每一类tf结合的DNA序列中存在特定类别的特征。结果:在本文中,我们应用稀疏贝叶斯学习算法来识别由不同类别的tf结合的DNA序列中的一小部分类别特异性特征;该算法同时学习一个真正的多类分类器,该分类器使用这些特征来预测识别特定DNA序列的TF的DNA结合域。我们在TRANSFAC中六个最大的类上训练我们的算法,总共包括587个tf。我们为这个训练集学习了一个六类分类器,它达到了87%的留一交叉验证准确率。我们还确定了顺式调控序列中对每一类TF具有高度特异性的特征,这对于如何为基序发现而对TF结合位点进行建模具有重要意义。
Motivation: A key goal in molecular biology is to understand the mechanisms by which a cell regulates the transcription of its genes. One important aspect of this transcriptional regulation is the binding of transcription factors (TFs) to their specific cis-regulatory counterparts on the DNA. TFs recognize and bind their DNA counterparts according to the structure of their DNA-binding domains (e.g. zinc finger, leucine zipper, homeodomain). The structure of these domains can be used as a basis for grouping TFs into classes. Although the structure of DNA-binding domains varies widely across TFs generally, the TFs within a particular class bind to DNA in a similar fashion, suggesting the existence of class-specific features in the DNA sequences bound by each class of TFs.Results: In this paper, we apply a sparse Bayesian learning algorithm to identify a small set of class-specific features in the DNA sequences bound by different classes of TFs; the algorithm simultaneously learns a true multi-class classifier that uses these features to predict the DNA-binding domain of the TF that recognizes a particular set of DNA sequences. We train our algorithm on the six largest classes in TRANSFAC, comprising a total of 587 TFs. We learn a six-class classifier for this training set that achieves 87% leave-one-out cross-validation accuracy. We also identify features within cis-regulatory sequences that are highly specific to each class of TF, which has significant implications for how TF binding sites should be modeled for the purpose of motif discovery.