MONONUCLEOTIDE THROUGH HEXANUCLEOTIDE COMPOSITION OF THE ESCHERICHIA-COLI GENOME - A MARKOV-CHAIN ANALYSIS

MONONUCLEOTIDE THROUGH HEXANUCLEOTIDE COMPOSITION OF THE ESCHERICHIA-COLI GENOME - A MARKOV-CHAIN ANALYSIS
复制标题

DOI:
10.1093/nar/15.6.2611
复制
发表时间:
1987-03-25
影响因子:
14.9
通讯作者:
IVARIE, R
IVARIE, R
中科院分区:
生物学2区
文献类型:
--
作者:
PHILLIPS, GJ;ARNOLD, J;IVARIE, R

文献摘要

被引文献

相似文献

对几种统计方法在预测大肠杆菌74,444 bp中二至六核苷酸观察频率的准确性进行了测试。coli DNA。总体而言,马尔可夫链是最准确的,而其他方法,包括基于潮汐频率的随机模型,则非常不准确。当从最高到最低丰度排序时,在E.科尔DNA高度不对称。所有有序丰度图具有宽的线性范围,包含在曲线的高端和低端急剧偏离的大多数低聚物。一般来说,马尔可夫链预测的值密切遵循有序丰度曲线的整体形状。推导出一个简单的公式,用该公式计算E.大肠杆菌基因组(或任何基因组)可以通过连续应用三阶马尔可夫链从三核苷酸和四核苷酸的嵌套组中相对准确地估计。该方程得出4,096个六核苷酸的预测频率与预期频率的平均比值为1.03±0.94。因此,该方法是六核苷酸位点之间的核苷酸长度的相对准确但不完美的预测器。使用四阶马尔可夫链和更大的数据集可以实现更高的准确性。寡核苷酸丰度的高度不对称性意味着E.大肠杆菌基因组全长4.2 × 106 bp,许多7-9 bp的相对短的序列非常罕见或缺失。
Several statistical methods were tested for accuracy in predicting observed frequencies of di- through hexanucleotides in 74,444 bp of E. coli DNA. A Markov chain was most accurate overall, whereas other methods, including a random model based on mononucleotide frequencies, were very inaccurate. When ranked highest to lowest abundance, the observed frequencies of oligonucleotides up to six bases in length in E. coll DNA were highly asymmetric. All ordered abundance plots had a wide linear range containing the majority of the oligomers which deviated sharply at the high and low ends of the curves. In general, values predicted by a Markov chain closely followed the overall shape of the ordered abundance curves. A simple equation was derived by which the frequency of any nucleotide longer than four bases in the E. coli genome (or any genome) can be relatively accurately estimated from the nested set of component tri- and tetranucleo-tides by serial application of a 3rd order Markov chain. The equation yielded a mean ratio of 1.03±0.94 for the observed-to-expected frequencies of the 4,096 hexanucleotides. Hence, the method is a relatively accurate but not perfect predictor of the length in nucleotides between hexanucleotide sites. Higher accuracy can be achieved using a 4th order Markov chain and larger data sets. The high asymmetry in oligonucleotide abundance neans that in the E. coli genome of 4.2 106bp many relatively short sequences of 7-9 bp are very rare or absent.