Duplication count distributions in DNA sequences.

Duplication count distributions in DNA sequences.
复制标题

DOI:
10.1103/physreve.78.061912
复制
发表时间:
2008-12
期刊:
Physical review. E, Statistical, nonlinear, and soft matter physics
影响因子:
--
通讯作者:
Yorke JA
Yorke JA
中科院分区:
其他
文献类型:
--
作者:
Sindi SS;Hunt BR;Yorke JA

文献摘要

被引文献

相似文献

我们通过研究足够长的序列来研究几个基因组中复杂重复DNA的定量特征,这些序列不太可能偶然重复。对于我们研究的每个基因组,我们确定相同拷贝的数量,即长度为40的每个序列的“复制数”,即每个“40-mer”。如果一个40-mer的重复数至少为2,我们就说它是“重复的”。我们主要关注“复杂”的40-mers,那些没有短的内部重复。我们发现我们可以将大多数复杂的重复40-mers分为两类:一类是它的拷贝紧密地聚集在一条染色体上,另一类是它的拷贝广泛地分布在多条染色体上。对于每个基因组和上述每个类别,我们计算N(c),即对于每个整数c,具有重复计数c的40-mers的数量。在每种情况下,我们观察到随着c从3增加到50或更高,N(c)呈幂律式衰减。特别是,我们发现N(c)的衰变比进化模型预测的要慢得多,在进化模型中,每个40-mer都有可能被复制。我们还分析了一个进化模型,它确实反映了N(c)的缓慢衰变。
We study quantitative features of complex repetitive DNA in several genomes by studying sequences that are sufficiently long that they are unlikely to have repeated by chance. For each genome we study, we determine the number of identical copies, the “duplication count,” of each sequence of length 40, that is of each “40-mer.” We say a 40-mer is “repeated” if its duplication count is at least 2. We focus mainly on “complex” 40-mers, those without short internal repetitions. We find that we can classify most of the complex repeated 40-mers into two categories: one category has its copies clustered closely together on one chromosome, the other has its copies distributed widely across multiple chromosomes. For each genome and each of the categories above, we compute N(c), the number of 40-mers that have duplication count c, for each integer c. In each case, we observe a power-law-like decay in N(c) as c increases from 3 to 50 or higher. In particular, we find that N(c) decays much more slowly than would be predicted by evolutionary models where each 40-mer is equally likely to be duplicated. We also analyze an evolutionary model that does reflect the slow decay of N(c).