Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.

Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.
复制标题

DOI:
--
复制
发表时间:
1996
期刊:
Computer applications in the biosciences : CABIOS
影响因子:
--
通讯作者:
K. Sjolander;K. Karplus;Michael Brown;R. Hughey;A. Krogh;1. S. Mian;D. Haussler
K. Sjolander;K. Karplus;Michael Brown;R. Hughey;A. Krogh;1. S. Mian;D. Haussler
中科院分区:
其他
文献类型:
--
作者:
K. Sjolander;K. Karplus;Michael Brown;R. Hughey;A. Krogh;1. S. Mian;D. Haussler

文献摘要

被引文献

相似文献

我们提出了一种将蛋白质多重比对中的信息压缩为氨基酸分布上狄利克雷密度的混合物的方法。狄利克雷混合密度被设计为与观察到的氨基酸频率相结合,以形成概况、隐马尔可夫模型或其他统计模型中每个位置的预期氨基酸概率的估计。这些估计赋予统计模型更大的泛化能力,以便模型可以更可靠地识别远亲家庭成员。本文纠正了之前发布的用于估计这些预期概率的公式,并包含狄利克雷混合公式的完整推导、优化混合物以匹配特定数据库的方法以及有效实施的建议。
We present a method for condensing the information in multiple alignments of proteins into a mixture of Dirichlet densities over amino acid distributions. Dirichlet mixture densities are designed to be combined with observed amino acid frequencies to form estimates of expected amino acid probabilities at each position in a profile, hidden Markov model or other statistical model. These estimates give a statistical model greater generalization capacity, so that remotely related family members can be more reliably recognized by the model. This paper corrects the previously published formula for estimating these expected probabilities, and contains complete derivations of the Dirichlet mixture formulas, methods for optimizing the mixtures to match particular databases, and suggestions for efficient implementation.