Accurate inference of transcription factor binding from DNA sequence and chromatin accessibility data

Accurate inference of transcription factor binding from DNA sequence and chromatin accessibility data
复制标题

DOI:
10.1101/gr.112623.110
复制
发表时间:
2011-03-01
期刊:
影响因子:
7
通讯作者:
Pritchard, Jonathan K.
Pritchard, Jonathan K.
中科院分区:
生物学1区
文献类型:
--
作者:
Pique-Regi, Roger;Degner, Jacob F.;Pritchard, Jonathan K.

文献摘要

被引文献

相似文献

准确的功能注释调控元件是理解全球基因调控的必要条件。在这里,我们报告了人类淋巴母细胞细胞系827,000个转录因子结合位点的全基因组图谱,其中包括239个已知转录因子结合基序的位置权重矩阵和49个新的序列基序。为了生成这张图谱,我们开发了一个概率框架,将细胞或组织特异性实验数据(如组蛋白修饰和dna酶I切割模式)与基因组信息(如基因注释和进化保护)结合起来。与经验ChIP-seq数据的比较表明,我们的方法是高度准确的,但具有优势,针对许多因素在一个单一的分析。我们预计这种方法将成为一种有价值的工具,用于各种细胞类型或组织在不同条件下的基因调控的全基因组研究。
Accurate functional annotation of regulatory elements is essential for understanding global gene regulation. Here, we report a genome-wide map of 827,000 transcription factor binding sites in human lymphoblastoid cell lines, which is comprised of sites corresponding to 239 position weight matrices of known transcription factor binding motifs, and 49 novel sequence motifs. To generate this map, we developed a probabilistic framework that integrates cell-or tissue-specific experimental data such as histone modifications and DNase I cleavage patterns with genomic information such as gene annotation and evolutionary conservation. Comparison to empirical ChIP-seq data suggests that our method is highly accurate yet has the advantage of targeting many factors in a single assay. We anticipate that this approach will be a valuable tool for genome-wide studies of gene regulation in a wide variety of cell types or tissues under diverse conditions.