Combining phylogenetic data with co-regulated genes to identify regulatory motifs

Combining phylogenetic data with co-regulated genes to identify regulatory motifs
复制标题

DOI:
10.1093/bioinformatics/btg329
复制
发表时间:
2003-12-12
期刊:
影响因子:
5.8
通讯作者:
Stormo, GD
Stormo, GD
中科院分区:
生物学3区
文献类型:
--
作者:
Wang, T;Stormo, GD

文献摘要

被引文献

相似文献

动机:在未对齐的DNA序列中发现调控基序仍然是计算生物学中的一个基本问题。已经开发了两类算法来从一组DNA序列中识别共同的基序。第一种方法可以称为“多基因,单物种”方法。它提出了一个简并基序嵌入在一些或所有其他无关的输入序列,并试图描述一个共识基序,并确定其出现。它通常用于通过实验方法鉴定的共调节基因。第二种方法可以称为“单基因,多物种”。它需要正向输入序列,并试图通过系统发育足迹来识别异常保守的区域。这两种方法都表现良好,但都有一些局限性。联合收割机结合不同基因之间的协同调节和同源基因之间的保守性来提高我们识别motif.Results的能力是很有吸引力的,基于我们小组以前建立的一致性算法,我们引入了一个新的算法称为PhyloCon(Phylogenetic Consensus),考虑到同源基因之间的保守性和物种内基因的协同调节。该算法首先将正向同源序列的保守区域比对成多个序列比对或谱,然后比较代表非正向同源序列的谱。在这些轮廓中,图案作为共同区域出现。在这里,我们提出了一种新的统计比较档案的DNA序列和贪婪的方法来寻找共同的子配置文件。我们证明了PhyloCon在合成和生物数据上都表现良好。
Motivation: Discovery of regulatory motifs in unaligned DNA sequences remains a fundamental problem in computational biology. Two categories of algorithms have been developed to identify common motifs from a set of DNA sequences. The first can be called a 'multiple genes, single species' approach. It proposes that a degenerate motif is embedded in some or all of the otherwise unrelated input sequences and tries to describe a consensus motif and identify its occurrences. It is often used for co-regulated genes identified through experimental approaches. The second approach can be called 'single gene, multiple species'. It requires orthologous input sequences and tries to identify unusually well conserved regions by phylogenetic footprinting. Both approaches perform well, but each has some limitations. It is tempting to combine the knowledge of co-regulation among different genes and conservation among orthologous genes to improve our ability to identify motifs.Results: Based on the Consensus algorithm previously established by our group, we introduce a new algorithm called PhyloCon (Phylogenetic Consensus) that takes into account both conservation among orthologous genes and co-regulation of genes within a species. This algorithm first aligns conserved regions of orthologous sequences into multiple sequence alignments, or profiles, then compares profiles representing non-orthologous sequences. Motifs emerge as common regions in these profiles. Here we present a novel statistic to compare profiles of DNA sequences and a greedy approach to search for common subprofiles. We demonstrate that PhyloCon performs well on both synthetic and biological data.