Application of experimentally verified transcription factor binding sites models for computational analysis of ChIP-Seq data.

Application of experimentally verified transcription factor binding sites models for computational analysis of ChIP-Seq data.
复制标题

DOI:
10.1186/1471-2164-15-80
复制
发表时间:
2014-01-29
期刊:
影响因子:
4.4
通讯作者:
Merkulova TI
Merkulova TI
中科院分区:
生物学2区
文献类型:
--
作者:
Levitsky VG;Kulakovskiy IV;Ershov NI;Oshchepkov DY;Makeev VJ;Hodgman TC;Merkulova TI

文献摘要

参考文献

被引文献

相似文献

ChIP-Seq广泛用于检测由转录因子(TF)直接在DNA结合位点(BS)或通过其他蛋白间接结合的基因组片段。目前,有许多软件工具实现了不同的方法来识别ChIP-Seq峰内的TFBS。然而,由于缺乏直接的实验验证,它们用于ChIP-Seq数据的解释通常很复杂,这使得难以设置阈值以避免识别太多的假阳性BS,并比较不同模型的实际性能。使用小鼠成年肝脏和人HepG 2细胞中FoxA 2结合位点的ChIP-Seq数据,我们比较了两种基本类别的四种计算模型的FoxA结合位点预测:基于实验证实的TFBS(oPWM和SiteGA)和从头基序发现(ChIPMunk和diChIPMunk)的现有训练集的模式匹配。为了正确选择模型的预测阈值,我们使用EMSA实验性地评估了64个预测的FoxA BS的亲和力,该EMSA允许安全地区分能够结合TF的序列。因此,我们在来自小鼠肝脏和人HepG 2细胞的ChIP-Seq基因座内鉴定出数千个可靠的FoxA BS。结果发现,传统的位置权重矩阵(PWM)模型的性能较差,误报率最高。相反,通过SiteGA和diChIPMunk/ChIPMunk模型的组合实现了最佳识别效率,在小鼠和人ChIP-Seq数据集的高达90%的基因座中正确识别FoxA BS。TF与预测位点对应的寡核苷酸结合的实验研究增加了ChIP-Seq数据分析中TFBS识别的计算方法的可靠性。关于ChIP-Seq数据解释,与更复杂的SiteGA和从头基序发现方法相比,基本PWM具有较差的TFBS识别质量。来自不同原理的模型的组合允许识别适当的TFBS。本文的在线版本(doi:10.1186/1471-2164-15-80)包含补充材料,可供授权用户使用。
ChIP-Seq is widely used to detect genomic segments bound by transcription factors (TF), either directly at DNA binding sites (BSs) or indirectly via other proteins. Currently, there are many software tools implementing different approaches to identify TFBSs within ChIP-Seq peaks. However, their use for the interpretation of ChIP-Seq data is usually complicated by the absence of direct experimental verification, making it difficult both to set a threshold to avoid recognition of too many false-positive BSs, and to compare the actual performance of different models. Using ChIP-Seq data for FoxA2 binding loci in mouse adult liver and human HepG2 cells we compared FoxA binding-site predictions for four computational models of two fundamental classes: pattern matching based on existing training set of experimentally confirmed TFBSs (oPWM and SiteGA) and de novo motif discovery (ChIPMunk and diChIPMunk). To properly select prediction thresholds for the models, we experimentally evaluated affinity of 64 predicted FoxA BSs using EMSA that allows safely distinguishing sequences able to bind TF. As a result we identified thousands of reliable FoxA BSs within ChIP-Seq loci from mouse liver and human HepG2 cells. It was found that the performance of conventional position weight matrix (PWM) models was inferior with the highest false positive rate. On the contrary, the best recognition efficiency was achieved by the combination of SiteGA & diChIPMunk/ChIPMunk models, properly identifying FoxA BSs in up to 90% of loci for both mouse and human ChIP-Seq datasets. The experimental study of TF binding to oligonucleotides corresponding to predicted sites increases the reliability of computational methods for TFBS-recognition in ChIP-Seq data analysis. Regarding ChIP-Seq data interpretation, basic PWMs have inferior TFBS recognition quality compared to the more sophisticated SiteGA and de novo motif discovery methods. A combination of models from different principles allowed identification of proper TFBSs. The online version of this article (doi:10.1186/1471-2164-15-80) contains supplementary material, which is available to authorized users.
DOI: 10.1371/journal.pbio.1001046
发表时间: 2011-04
期刊: PLoS biology
影响因子: 9.8
作者:
ENCODE Project Consortium
通讯作者: ENCODE Project Consortium
DOI: 10.1016/j.gde.2010.06.005
发表时间: 2010-10
影响因子: 4
作者:
Kaestner, Klaus H.
通讯作者: Kaestner, Klaus H.
蛋白质结合阵列数据的新计算分析确定了胰腺中NKX2.2的直接靶标。
DOI: 10.1186/1471-2105-12-62
发表时间: 2011-02-25
期刊: BMC bioinformatics
影响因子: 3
作者:
Hill JT;Anderson KR;Mastracci TL;Kaestner KH;Sussel L
通讯作者: Sussel L
DOI: 10.1128/mcb.18.11.6305
发表时间: 1998-11-01
影响因子: 5.3
作者:
Christoffels, VM;Grange, T;Lamers, WH
通讯作者: Lamers, WH
DOI: 10.1186/1471-2105-7-279
发表时间: 2006-06-02
期刊: BMC bioinformatics
影响因子: 3
作者:
Huang W;Umbach DM;Ohler U;Li L
通讯作者: Li L