Motto: Representing Motifs in Consensus Sequences with Minimum Information Loss.

Motto: Representing Motifs in Consensus Sequences with Minimum Information Loss.
复制标题

DOI:
10.1534/genetics.120.303597
复制
发表时间:
2020-10
期刊:
影响因子:
3.3
通讯作者:
Wang W
Wang W
中科院分区:
生物学2区
文献类型:
--
作者:
Wang M;Wang D;Zhang K;Ngo V;Fan S;Wang W

文献摘要

相似文献

序列分析经常需要对基序的直观理解和方便的表示。典型地,基序被表示为位置权重矩阵(PWM)并且使用序列标志可视化。然而,在许多情况下,为了解释基序信息或搜索基序匹配,通过通配符样式的共有序列(例如[GC][AT]GATAAG[GAC])来表示基序是紧凑且足够的。基于互信息理论和Jensen-Shannon散度,我们提出了一个数学框架,以最大限度地减少PWM到共识序列的信息损失。我们将这种表示命名为序列座右铭,并实现了一个有效的算法,灵活的选项,用于将基序PWM转换为座右铭从核苷酸,氨基酸和自定义字符。我们表明,这种表示提供了一个简单而有效的方法来确定1156个共同的转录因子(TF)在人类基因组中的结合位点。该方法的有效性进行了基准比较序列匹配发现的座右铭与PWM扫描结果发现FIMO。平均而言,我们的方法达到了0.81的精确度-召回率曲线下的面积,显著(P值< 0.01)优于所有现有的方法,包括最大位置权重,Cavener方法和最小均方误差。我们相信,这种表示提供了一个主题的蒸馏总结,以及统计的理由。
Sequence analysis frequently requires intuitive understanding and convenient representation of motifs. Typically, motifs are represented as position weight matrices (PWMs) and visualized using sequence logos. However, in many scenarios, in order to interpret the motif information or search for motif matches, it is compact and sufficient to represent motifs by wildcard-style consensus sequences (such as [GC][AT]GATAAG[GAC]). Based on mutual information theory and Jensen-Shannon divergence, we propose a mathematical framework to minimize the information loss in converting PWMs to consensus sequences. We name this representation as sequence Motto and have implemented an efficient algorithm with flexible options for converting motif PWMs into Motto from nucleotides, amino acids, and customized characters. We show that this representation provides a simple and efficient way to identify the binding sites of 1156 common transcription factors (TFs) in the human genome. The effectiveness of the method was benchmarked by comparing sequence matches found by Motto with PWM scanning results found by FIMO. On average, our method achieves a 0.81 area under the precision-recall curve, significantly (P-value < 0.01) outperforming all existing methods, including maximal positional weight, Cavener’s method, and minimal mean square error. We believe this representation provides a distilled summary of a motif, as well as the statistical justification.