Probabilistic and statistical properties of words: An overview

Probabilistic and statistical properties of words: An overview
复制标题

DOI:
10.1089/10665270050081360
复制
发表时间:
2000-02-01
影响因子:
1.7
通讯作者:
Waterman, MS
Waterman, MS
中科院分区:
生物学4区
文献类型:
--
作者:
Reinert, G;Schbath, S;Waterman, MS

文献摘要

被引文献

相似文献

在下文中,概述了在生物序列分析中出现的单词的统计和概率特性。计数的发生,计数的团块,和更新计数的区别,以及正常的近似,泊松过程近似和复合泊松近似的精确分布。在这里,一个序列被建模为一个平稳遍历马尔可夫链,用于确定适当的顺序的马尔可夫链的测试进行说明。收敛性结果考虑了马尔可夫转移概率估计的误差,主要工具是矩母函数、鞅、Stein方法和Chen-Stein方法。类似的结果给出了多个模式的出现,并作为一个例子,从SBH芯片数据的序列的唯一可恢复性的问题进行了讨论,特别强调在于解开字出现之间的复杂的依赖结构,由于自重叠,以及由于字之间的重叠。结果可用于推导近似和保守的置信区间。
In the following, an overview is given on statistical and probabilistic properties of words, as occurring in the analysis of biological sequences. Counts of occurrence, counts of clumps, and renewal counts are distinguished, and exact distributions as well as normal approximations, Poisson process approximations, and compound Poisson approximations are derived. Here, a sequence is modelled as a stationary ergodic Markov chain; a test for determining the appropriate order of the Markov chain is described. The convergence results take the error made by estimating the Markovian transition probabilities into account, The main tools involved are moment generating functions, martingales, Stein's method, and the Chen-Stein method. Similar results are given for occurrences of multiple patterns, and, as an example, the problem of unique recoverability of a sequence from SBH chip data is discussed, Special emphasis lies on disentangling the complicated dependence structure between word occurrences, due to self-overlap as well as due to overlap between words. The results can be used to derive approximate, and conservative, confidence intervals for tests.