DNA familial binding profiles made easy: comparison of various motif alignment and clustering strategies.

DNA familial binding profiles made easy: comparison of various motif alignment and clustering strategies.
复制标题

DOI:
10.1371/journal.pcbi.0030061
复制
发表时间:
2007-03-30
影响因子:
4.3
通讯作者:
Benos, Panayiotis V.
Benos, Panayiotis V.
中科院分区:
生物学2区
文献类型:
--
作者:
Mahony, Shaun;Auron, Philip E.;Benos, Panayiotis V.

文献摘要

参考文献

被引文献

相似文献

转录因子(TF)蛋白以高特异性识别少量DNA序列并控制邻近基因的表达。 TF 结合偏好的演变已成为许多近期研究的主题,其中引入了广义结合谱并用于改进新靶位点的预测。通用配置文件是通过对齐和合并相关 TF 的各个配置文件来生成的。然而,用于比较结合概况的距离度量和比对算法尚未得到充分探索或优化。因此,结合谱取决于 TF 结构信息,有时可能会忽略亚家族之间的重要区别。预测与给定 DNA 模式结合的蛋白质的身份或结构类别将增强微阵列和 ChIP 芯片数据的分析,其中经常预测通常未知的 TF 的多个假定目标。评估各种比较指标和对齐算法(总共 105 种组合)。我们发现,在检测真核 DNA 基序相似性方面,局部比对通常比全局比对更好,尤其是与距离平方和或 Pearson 相关系数比较指标相结合时。此外,还测试了结合谱的多重比对策略和树构建方法在构建广义结合模型方面的效率。开发了一种自动确定最佳簇数的新方法,并将其应用于构建一组新的家族结合谱,从而提高了 TF 分类的准确性。开发了一个软件工具 STAMP 来托管所有经过测试的方法并使它们公开可用。这项工作提供了高质量的家族结合图谱参考集和第一个用于 DNA 图谱分析的综合平台。检测DNA基序之间的相似性是转录调控比较研究的关键步骤,这里介绍的工作将为未来转录建模研究的工具和方法开发奠定基础。转录因子是基因表达的主要调节因子。它们通常识别基因启动子中的短 DNA 序列,并随后改变其​​转录率。众所周知,结构相关的转录因子通常识别相似的 DNA 结合模式(或基序)。这些图案的比较不仅可以深入了解它们所经历的进化过程,而且还具有许多重要的实际应用。例如,发现“相似”的基序可以组合起来形成广义图谱,这可以用来提高我们预测共表达基因启动子中新DNA信号的能力,从而有助于更准确地绘制基因调控网络。然而,迄今为止,还没有可以有效分析 DNA 基序的综合平台。此外,用于分配 DNA 基序之间相似性的方法的效率尚未经过彻底测试。本文通过评估应用于 DNA 基序的可用比较策略并生成改进的家族概况数据集,朝着这一目标迈出了重要的第一步。
Transcription factor (TF) proteins recognize a small number of DNA sequences with high specificity and control the expression of neighbouring genes. The evolution of TF binding preference has been the subject of a number of recent studies, in which generalized binding profiles have been introduced and used to improve the prediction of new target sites. Generalized profiles are generated by aligning and merging the individual profiles of related TFs. However, the distance metrics and alignment algorithms used to compare the binding profiles have not yet been fully explored or optimized. As a result, binding profiles depend on TF structural information and sometimes may ignore important distinctions between subfamilies. Prediction of the identity or the structural class of a protein that binds to a given DNA pattern will enhance the analysis of microarray and ChIP–chip data where frequently multiple putative targets of usually unknown TFs are predicted. Various comparison metrics and alignment algorithms are evaluated (a total of 105 combinations). We find that local alignments are generally better than global alignments at detecting eukaryotic DNA motif similarities, especially when combined with the sum of squared distances or Pearson's correlation coefficient comparison metrics. In addition, multiple-alignment strategies for binding profiles and tree-building methods are tested for their efficiency in constructing generalized binding models. A new method for automatic determination of the optimal number of clusters is developed and applied in the construction of a new set of familial binding profiles which improves upon TF classification accuracy. A software tool, STAMP, is developed to host all tested methods and make them publicly available. This work provides a high quality reference set of familial binding profiles and the first comprehensive platform for analysis of DNA profiles. Detecting similarities between DNA motifs is a key step in the comparative study of transcriptional regulation, and the work presented here will form the basis for tool and method development for future transcriptional modeling studies. Transcription factors are primary regulators of gene expression. They usually recognize short DNA sequences in gene promoters and subsequently alter their transcription rate. It is known that structurally related transcription factors often recognize similar DNA-binding patterns (or motifs). Comparison of these motifs not only provides insights into the evolutionary process they undergo, but it also has many important practical applications. For example, motifs that are found to be “similar” can be combined to form generalized profiles, which can be used to improve our ability to predict novel DNA signals in the promoters of co-expressed genes, and thus facilitate a more accurate mapping of gene-regulatory networks. However, to date there is no comprehensive platform that will allow for an efficient analysis of DNA motifs. Furthermore, the efficiency of the methods used to assign similarity between DNA motifs has not been thoroughly tested. This paper takes an important first step towards this goal by evaluating available comparison strategies as applied to DNA motifs and by generating an improved familial profile dataset.
DOI: 10.1186/1471-2105-6-237
发表时间: 2005-09-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Kielbasa SM;Gonze D;Herzel H
通讯作者: Herzel H
DOI: 10.1093/oxfordjournals.molbev.a040454
发表时间: 1987-07-01
影响因子: 10.7
作者:
SAITOU, N;NEI, M
通讯作者: NEI, M
DOI: 10.1093/bioinformatics/bti473
发表时间: 2005-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cartharius, K;Frech, K;Werner, T
通讯作者: Werner, T
DOI: 10.1093/bioinformatics/bti815
发表时间: 2006-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
MacIsaac, KD;Gordon, DB;Fraenkel, E
通讯作者: Fraenkel, E
DOI: 10.1093/bioinformatics/bti731
发表时间: 2006-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Narlikar, L;Hartemink, AJ
通讯作者: Hartemink, AJ