Modeling dependencies in protein-DNA binding sites

Modeling dependencies in protein-DNA binding sites
复制标题

DOI:
10.1145/640075.640079
复制
发表时间:
2003-04
期刊:
--
影响因子:
--
通讯作者:
Yoseph Barash;G. Elidan;N. Friedman;Tommy Kaplan
Yoseph Barash;G. Elidan;N. Friedman;Tommy Kaplan
中科院分区:
其他
文献类型:
--
作者:
Yoseph Barash;G. Elidan;N. Friedman;Tommy Kaplan

文献摘要

被引文献

相似文献

全基因组序列和高通量基因组测定的可用性为转录调控的计算机分析打开了大门。这包括用于发现和表征DNA结合蛋白(如转录因子)的结合位点的方法。转录因子结合位点的常见表示是位置特异性得分矩阵(PSSM)。该表示法强烈假设结合位点位置彼此独立。在这项工作中,我们探索贝叶斯网络表示的结合位点,提供不同的权衡复杂性(参数的数量)和丰富的位置之间的依赖关系。我们开发了正式的机器,从数据中学习这样的模型,并估计推定的结合位点的统计意义。然后,我们评估这些更丰富的表征结合位点图案和预测其基因组位置的后果。我们表明,这些更丰富的表示改善了PSSM模型在这两个任务。
The availability of whole genome sequences and high-throughput genomic assays opens the door for in silico analysis of transcription regulation. This includes methods for discovering and characterizing the binding sites of DNA-binding proteins, such as transcription factors. A common representation of transcription factor binding sites is a position specific score matrix (PSSM). This representation makes the strong assumption that binding site positions are independent of each other. In this work, we explore Bayesian network representations of binding sites that provide different tradeoffs between complexity (number of parameters) and the richness of dependencies between positions. We develop the formal machinery for learning such models from data and for estimating the statistical significance of putative binding sites. We then evaluate the ramifications of these richer representations in characterizing binding site motifs and predicting their genomic locations. We show that these richer representations improve over the PSSM model in both tasks.