A statistical method for alignment-free comparison of regulatory sequences

A statistical method for alignment-free comparison of regulatory sequences
复制标题

DOI:
10.1093/bioinformatics/btm211
复制
发表时间:
2007-07-01
期刊:
影响因子:
5.8
通讯作者:
Sinha, Saurabh
Sinha, Saurabh
中科院分区:
生物学3区
文献类型:
--
作者:
Kantorovitz, Miriam R.;Robinson, Gene E.;Sinha, Saurabh

文献摘要

被引文献

相似文献

动机:两个生物序列的相似性传统上是在公认的比对框架内评估的。在这里,我们专注于确定顺式调控序列之间的功能关系,是非正交或大大分歧的任务。在此机制中需要“无比对”的序列相似性测量。结果:我们研究了用于无比对序列比较的新评分(称为D2 z评分)的使用。它基于比较两个序列中所有固定长度字的频率。分数的一个重要的新颖特征是,它在从任意背景分布绘制的序列对之间是可比较的。我们提出了一种方法,在计算D2z分数的时间复杂度上给出了二次改进。然后,我们评估几个组织特异性家族的顺式调控模块(在果蝇和人类)的分数。新的评分是非常成功的区分功能相关的调控序列无关的序列对。的D2z分数的性能进行比较,其他五个无障碍的相似性措施,并显示出一贯优于所有这些措施的上级。
Motivation: The similarity of two biological sequences has traditionally been assessed within the well-established framework of alignment. Here we focus on the task of identifying functional relationships between cis-regulatory sequences that are non-orthologous or greatly diverged. 'Alignment-free' measures of sequence similarity are required in this regime.Results: We investigate the use of a new score for alignment-free sequence comparison, called the D2z score. It is based on comparing the frequencies of all fixed-length words in the two sequences. An important, novel feature of the score is that it is comparable across sequence pairs drawn from arbitrary background distributions. We present a method that gives quadratic improvement in the time complexity of calculating the D2z score, over the naive method. We then evaluate the score on several tissue-specific families of cis-regulatory modules ( in Drosophila and human). The new score is highly successful in discriminating functionally related regulatory sequences from unrelated sequence pairs. The performance of the D2z score is compared to five other alignment-free similarity measures, and shown to be consistently superior to all of these measures.