Using sequence-specific chemical and structural properties of DNA to predict transcription factor binding sites.

Using sequence-specific chemical and structural properties of DNA to predict transcription factor binding sites.
复制标题

DOI:
10.1371/journal.pcbi.1001007
复制
发表时间:
2010-11-18
影响因子:
4.3
通讯作者:
Mu F
Mu F
中科院分区:
生物学2区
文献类型:
--
作者:
Bauer AL;Hlavacek WS;Unkefer PJ;Mu F

文献摘要

参考文献

被引文献

相似文献

An important step in understanding gene regulation is to identify the DNA binding sites recognized by each transcription factor (TF). Conventional approaches to prediction of TF binding sites involve the definition of consensus sequences or position-specific weight matrices and rely on statistical analysis of DNA sequences of known binding sites. Here, we present a method called SiteSleuth in which DNA structure prediction, computational chemistry, and machine learning are applied to develop models for TF binding sites. In this approach, binary classifiers are trained to discriminate between true and false binding sites based on the sequence-specific chemical and structural features of DNA. These features are determined via molecular dynamics calculations in which we consider each base in different local neighborhoods. For each of 54 TFs in Escherichia coli, for which at least five DNA binding sites are documented in RegulonDB, the TF binding sites and portions of the non-coding genome sequence are mapped to feature vectors and used in training. According to cross-validation analysis and a comparison of computational predictions against ChIP-chip data available for the TF Fis, SiteSleuth outperforms three conventional approaches: Match, MATRIX SEARCH, and the method of Berg and von Hippel. SiteSleuth also outperforms QPMEME, a method similar to SiteSleuth in that it involves a learning algorithm. The main advantage of SiteSleuth is a lower false positive rate. An important step in characterizing the genetic regulatory network of a cell is to identify the DNA binding sites recognized by each transcription factor (TF) protein encoded in the genome. Current computational approaches to TF binding site prediction rely exclusively on DNA sequence analysis. In this manuscript, we present a novel method called SiteSleuth, in which classifiers are trained to discriminate between true and false binding sites based on the sequence-specific chemical and structural features of DNA. According to cross-validation analysis and a comparison of computational predictions against ChIP-chip data available for the TF Fis, SiteSleuth predicts fewer estimated false positives than any of four other methods considered. A better understanding of gene regulation, which plays a central role in cellular responses to environmental changes, is a key to manipulating cellular behavior for a variety of useful purposes, as in metabolic engineering applications.
DOI: 10.1038/nprot.2008.195
发表时间: 2009
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Berger, Michael F.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1126/science.1131007
发表时间: 2007-01-12
期刊: SCIENCE
影响因子: 56.9
作者:
Maerkl, Sebastian J.;Quake, Stephen R.
通讯作者: Quake, Stephen R.
DOI: 10.1002/0471143030.cb1707s23
发表时间: 2004-09-01
影响因子: --
作者:
Aparicio, Oscar;Geisberg, Joseph V;Struhl, Kevin
通讯作者: Struhl, Kevin
DOI: 10.1021/jm00145a002
发表时间: 1985-01-01
影响因子: 7.3
作者:
GOODFORD, PJ
通讯作者: GOODFORD, PJ
DOI: 10.1016/s0006-3495(92)81649-1
发表时间: 1992-09-01
影响因子: 3.4
作者:
BERMAN, HM;OLSON, WK;SCHNEIDER, B
通讯作者: SCHNEIDER, B