UNSUPERVISED LEARNING OF MULTIPLE MOTIFS IN BIOPOLYMERS USING EXPECTATION MAXIMIZATION

UNSUPERVISED LEARNING OF MULTIPLE MOTIFS IN BIOPOLYMERS USING EXPECTATION MAXIMIZATION
复制标题

DOI:
10.1007/bf00993379
复制
发表时间:
1995-10-01
期刊:
影响因子:
7.5
通讯作者:
ELKAN, C
ELKAN, C
中科院分区:
计算机科学3区
文献类型:
--
作者:
BAILEY, TL;ELKAN, C

文献摘要

被引文献

相似文献

MEME算法扩展了期望最大化(EM)算法,用于识别未比对的生物聚合物序列中的基序。MEME的目的是在一组生物聚合物序列中发现新的基序,对于可能存在的任何基序事先知之甚少或一无所知。MEME的创新扩展了可使用EM解决的问题范围,并增加了找到良好解决方案的机会。首先,将实际出现在生物聚合物序列中的子序列用作EM算法的起始点,以增加找到全局最优基序的概率。其次,去除了每个序列恰好包含共享基序的一次出现这一假设。这允许基序在任何序列中多次出现,并允许算法忽略没有共享基序出现的序列,提高了其对噪声数据的抗性。第三,纳入了一种在发现共享基序后以概率方式擦除它们的方法,以便在同一组序列中可以发现几个不同的基序,无论是不同基序出现在不同序列中,还是单个序列可能包含多个基序时。实验表明,MEME可以从一组包含一个或两个位点的序列中发现CRP和LexA结合位点,并且MEME可以在一组大肠杆菌序列中发现 -10和 -35启动子区域。
The MEME algorithm extends the expectation maximization (EM) algorithm for identifying motifs in unaligned biopolymer sequences. The aim of MEME is to discover new motifs in a set of biopolymer sequences where little or nothing is known in advance about any motifs that may be present. MEME innovations expand the range of problems which can be solved using EM and increase the chance of finding good solutions. First, subsequences which actually occur in the biopolymer sequences are used as starting points for the EM algorithm to increase the probability of finding globally optimal motifs. Second, the assumption that each sequence contains exactly one occurrence of the shared motif is removed. This allows multiple appearances of a motif to occur in any sequence and permits the algorithm to ignore sequences with no appearance of the shared motif, increasing its resistance to noisy data. Third, a method for probabilistically erasing shared motifs after they are found is incorporated so that several distinct motifs can be found in the same set of sequences, both when different motifs appear in different sequences and when a single sequence may contain multiple motifs. Experiments show that MEME can discover both the CRP and LexA binding sites from a set of sequences which contain one or both sites, and that MEME can discover both the -10 and -35 promoter regions in a set of E. coli sequences.