Integrating multiple evidence sources to predict transcription factor binding in the human genome

Integrating multiple evidence sources to predict transcription factor binding in the human genome
复制标题

DOI:
10.1101/gr.096305.109
复制
发表时间:
2010-04-01
期刊:
影响因子:
7
通讯作者:
Bar-Joseph, Ziv
Bar-Joseph, Ziv
中科院分区:
生物学1区
文献类型:
--
作者:
Ernst, Jason;Plasterer, Heather L.;Bar-Joseph, Ziv

文献摘要

被引文献

相似文献

关于许多转录因子的结合偏好的信息是已知的,并且其特征在于序列结合基序。然而,基于其基序确定转录因子结合的基因组区域是一个具有挑战性的问题,特别是在具有大基因组的物种中,因为通常存在许多包含与基序匹配但未结合的序列。基于序列保守性或相对于转录起始位点的位置的几个规则已经被提出来帮助区分真正的结合位点和随机的结合位点。其他证据来源也可能为这项任务提供信息。我们开发了一种使用逻辑回归分类器整合多个证据源的方法。我们的方法分为两步。首先,我们基于大量证据特征来推断量化所有位置转录因子结合的一般结合偏好的分数,而不使用任何基序特定信息。然后,我们将此一般结合偏好得分与特定转录因子的基序信息相结合,以提高对因子结合区域的预测。使用交叉验证和新的实验数据,我们表明,令人惊讶的是,一般的结合偏好可以高度预测的真实位置的转录因子结合,即使没有结合基序使用。当结合基序信息,我们的方法优于以前的方法预测真正的结合的位置。
Information about the binding preferences of many transcription factors is known and characterized by a sequence binding motif. However, determining regions of the genome in which a transcription factor binds based on its motif is a challenging problem, particularly in species with large genomes, since there are often many sequences containing matches to the motif but are not bound. Several rules based on sequence conservation or location, relative to a transcription start site, have been proposed to help differentiate true binding sites from random ones. Other evidence sources may also be informative for this task. We developed a method for integrating multiple evidence sources using logistic regression classifiers. Our method works in two steps. First, we infer a score quantifying the general binding preferences of transcription factor binding at all locations based on a large set of evidence features, without using any motif specific information. Then, we combined this general binding preference score with motif information for specific transcription factors to improve prediction of regions bound by the factor. Using cross-validation and new experimental data we show that, surprisingly, the general binding preference can be highly predictive of true locations of transcription factor binding even when no binding motif is used. When combined with motif information our method outperforms previous methods for predicting locations of true binding.