Sequence analysis by additive scales:: DNA structure for sequences and repeats of all lengths

Sequence analysis by additive scales:: DNA structure for sequences and repeats of all lengths
复制标题

DOI:
10.1093/bioinformatics/16.10.865
复制
发表时间:
2000-10-01
期刊:
影响因子:
5.8
通讯作者:
Baisnée, PF
Baisnée, PF
中科院分区:
生物学3区
文献类型:
--
作者:
Baldi, P;Baisnée, PF

文献摘要

被引文献

相似文献

动机:DNA结构在多种生物过程中起着重要作用。人们已经提出了不同的双核苷酸和三核苷酸尺度来捕捉DNA结构的各个方面,包括碱基堆叠能、螺旋桨扭转角、蛋白质的可变形性、可弯曲性和位置偏好,然而,仍然缺乏一个用于计算分析和预测DNA结构的通用框架。这样的框架应该特别解决以下问题:(1)构造具有外部属性的序列;(2)根据给定的基因组背景对序列进行定量评价;(3)从基因组数据库中自动提取极端序列和基因图谱;(4)随序列长度N增加的分布及渐近性态;(5)完整的尺度间相关性分析。结果:我们开发了一个基于附加尺度,结构或其他解决所有这些问题的序列分析的一般框架。我们展示了如何构建极值序列和校准自动基因组和数据库提取的分数。随着N的增加,分布迅速收敛到正态,尺度之间的两两相关性依赖于背景分布和序列长度,并迅速收敛到一个可分析预测的渐近值。对于双核苷酸和三核苷酸尺度,在大约10-15 bp的特征窗口长度上获得正常行为和渐近相关值。在均匀的背景分布下,经验推导的尺度之间的两两相关性保持相对较小,并且在所有长度上都大致恒定,除了螺旋桨捻度和蛋白质变形性呈正相关。二核苷酸碱基堆积(如螺旋扭和蛋白质变形能力)与at含量随长度的增加呈负相关。该框架适用于各种DNA串联重复序列的分析。我们导出了计算所有长度上重复单元类数目的精确表达式。串联重复可能是由多种不同的机制造成的,其中一部分可能取决于以极端结构特征为特征的谱。
Motivation: DNA structure plays an important role in a variety of biological processes. Different di- and trinucleotide scales have been proposed to capture various aspects of DNA structure including base stacking energy, propeller twist angle, protein deformability, bendability, and position preference, Yet, a general framework for the computational analysis and prediction of DNA structure is still lacking. Such a framework Should in particular address the following issues: (1) construction of sequences with external properties; (2) quantitative evaluation of sequences with respect to a given genomic background; (3) automatic extraction of extremal sequences and profiles from genomic databases; (4) distribution and asymptotic behavior as the length N of the sequences increases; and (5) complete analysis of correlations between scales.Results: We develop a general framework for sequence analysis based on additive scales, structural or other that addresses all these issues. WE show how to construct extremal sequences and calibrate scores for automatic genomic and database extraction. We show that distributions rapidly converge to normality as N increases, Pairwise correlations between scales depend both on background distribution and sequence length and rapidly converge to an analytically predictable asymptotic value. For di- and tri-nucleotide scales, normal behavior and asymptotic correlation values are attained over a characteristic window length of about 10-15 bp. With a uniform background distribution, pairwise correlations between empirically-derived scales remain relatively small and roughly constant at all lengths, except for propeller twist and protein deformability which are positively correlated There is a positive (resp. negative) correlation between dinucleotide base stacking (resp, propeller twist and protein deformability) and AT-content that increases in magnitude with length. The framework is applied to the analysis of various DNA tandem repeats. We derive exact expressions for counting the number of repeat unit classes at all lengths. Tandem repeats are likely to result from a variety of different mechanisms, a fraction of which is likely to depend on profiles characterized by extreme structural features.