Fitting a Mixture Model By Expectation Maximization To Discover Motifs In Biopolymer

Fitting a Mixture Model By Expectation Maximization To Discover Motifs In Biopolymer
复制标题

DOI:
--
复制
发表时间:
1994
期刊:
Proceedings. International Conference on Intelligent Systems for Molecular Biology
影响因子:
--
通讯作者:
T. Bailey;C. Elkan
T. Bailey;C. Elkan
中科院分区:
其他
文献类型:
--
作者:
T. Bailey;C. Elkan

文献摘要

被引文献

相似文献

本文所描述的算法发现一个或多个模体的DNA或蛋白质序列的集合中使用的技术的期望最大化,以适应一个两个组件的有限混合模型的序列集。通过将混合模型拟合到数据,概率性地擦除由此发现的基序的出现,并重复该过程以找到连续的基序,来找到多个基序。该算法只需要一组未对齐的序列和一个指定图案宽度的数字作为输入。它返回每个主题的模型和阈值,它们可以一起用作贝叶斯最优分类器,用于在其他数据库中搜索主题的出现。该算法估计每个基序在数据集中的每个序列中出现多少次,并输出基序出现的对齐。该算法能够发现几个不同的图案在一个单一的数据集中出现的不同数量。
The algorithm described in this paper discovers one or more motifs in a collection of DNA or protein sequences by using the technique of expectation maximization to fit a two-component finite mixture model to the set of sequences. Multiple motifs are found by fitting a mixture model to the data, probabilistically erasing the occurrences of the motif thus found, and repeating the process to find successive motifs. The algorithm requires only a set of unaligned sequences and a number specifying the width of the motifs as input. It returns a model of each motif and a threshold which together can be used as a Bayes-optimal classifier for searching for occurrences of the motif in other databases. The algorithm estimates how many times each motif occurs in each sequence in the dataset and outputs an alignment of the occurrences of the motif. The algorithm is capable of discovering several different motifs with differing numbers of occurrences in a single dataset.