Nonoverlapping clusters: approximate distribution and application to molecular biology.

Nonoverlapping clusters: approximate distribution and application to molecular biology.
复制标题

非重叠簇:近似分布及其在分子生物学中的应用。

DOI:
10.1111/j.0006-341x.2001.00420.x
复制
发表时间:
2001
期刊:
Biometrics.
影响因子:
--
通讯作者:
Bishop,D
Bishop,D
中科院分区:
--
文献类型:
--
作者:
Su,X;Wallenstein,S;Bishop,D

文献摘要

相似文献

开发了一种筛选基因组序列数据以识别基因调控区的方法。这种方法是基于确定假定的转录因子结合位点是否比人们偶然预期的更大程度地聚集在一起。给定事件发生在宽度L(Lbase对)的区间上,ANR:WCLUSTER被定义为包含在长度为L的窗口内的所有R+1个连续事件。在事件的位置具有均匀分布的模型下,给出了非重叠事件数分布的精确且易于计算的近似公式。仿真结果表明,这些近似方法比现有的方法具有更高的精度。该近似被应用于检测基因组DNA序列中的红系特异性调节区,首先在人工情况下指定优先顺序,然后作为探索性方法的一部分。
An approach is developed for the screening of genomic sequence data to identify gene regulatory regions. This approach is based on deciding if putative transcription factor binding sites are clustered together to a greater extent than one would expect by chance. Givennevents occurring on an interval of widthL(Lbase pairs), anr:wcluster is defined asr+ 1 consecutive events all contained within a window of lengthwL.Accurate and easily computable approximations are derived for the distribution of the number of nonoverlappingr:wclusters under the model that the positions of thenevents have a uniform distribution. Simulations demonstrate that these approximations have greater accuracy than existing methods. The approximation is applied to detect erythroid‐specific regulatory regions in genomic DNA sequences, first in an artificial case whereris specifieda prioriand then as part of an exploratory approach.