CUBIC: identification of regulatory binding sites through data clustering.

CUBIC: identification of regulatory binding sites through data clustering.
复制标题

DOI:
10.1142/s0219720003000162
复制
发表时间:
2003-04-01
影响因子:
1
通讯作者:
Xu, Ying
Xu, Ying
中科院分区:
生物学4区
文献类型:
--
作者:
Olman, Victor;Xu, Dong;Xu, Ying

文献摘要

被引文献

相似文献

转录因子结合位点是基因上游区域的短片段,转录因子与其结合以调节基因转录为 mRNA。尽管已经投入了大量的精力来研究转录因子结合位点的计算识别仍然是一个未解决的具有挑战性的问题。我们最近开发了一种新技术,用于从一组基因上游区域识别结合位点,这些区域可能是转录共同调节的,因此可能共享相似的转录因子结合位点。通过利用此类结合位点的两个关键特征(即它们的高序列相似性和与其他序列片段相比相对较高的频率),我们将该问题表述为簇识别问题。即从噪声背景中识别并提取数据簇。虽然经典的数据聚类问题(将数据集划分为具有共同或相似特征的簇)已被广泛研究,但还没有通用的算法来解决从噪声背景中识别数据簇的问题。在本文中,我们提出了一种解决此类问题的新颖算法。我们已经证明,根据我们的定义,可以通过在线性序列中搜索具有特殊属性的子串来严格有效地解决簇识别问题。我们还开发了一种评估每个已识别聚类的统计显着性的方法,可用于排除意外的数据聚类。我们将聚类识别算法和统计显着性分析方法实现为计算机软件CUBIC。对 CUBIC 进行了广泛的测试。我们在这里介绍了 CUBIC 在具有挑战性的结合位点识别案例中的一些应用。
Transcription factor binding sites are short fragments in the upstream regions of genes, to which transcription factors bind to regulate the transcription of genes into mRNA. Computational identification of transcription factor binding sites remains an unsolved challenging problem though a great amount of effort has been put into the study of this problem. We have recently developed a novel technique for identification of binding sites from a set of upstream regions of genes, that could possibly be transcriptionally co-regulated and hence might share similar transcription factor binding sites. By utilizing two key features of such binding sites (i.e. their high sequence similarities and their relatively high frequencies compared to other sequence fragments), we have formulated this problem as a cluster identification problem. That is to identify and extract data clusters from a noisy background. While the classical data clustering problem (partitioning a data set into clusters sharing common or similar features) has been extensively studied, there is no general algorithm for solving the problem of identifying data clusters from a noisy background. In this paper, we present a novel algorithm for solving such a problem. We have proved that a cluster identification problem, under our definition, can be rigorously and efficiently solved through searching for substrings with special properties in a linear sequence. We have also developed a method for assessing the statistical significance of each identified cluster, which can be used to rule out accidental data clusters. We have implemented the cluster identification algorithm and the statistical significance analysis method as a computer software CUBIC. Extensive testing on CUBIC has been carried out. We present here a few applications of CUBIC on challenging cases of binding site identification.