Information content of individual genetic sequences

Information content of individual genetic sequences
复制标题

DOI:
10.1006/jtbi.1997.0540
复制
发表时间:
1997-12-21
影响因子:
2
通讯作者:
Schneider, TD
Schneider, TD
中科院分区:
生物学4区
文献类型:
--
作者:
Schneider, TD

文献摘要

被引文献

相似文献

具有共同功能的相关基因序列可以通过香农信息测量来描述并通过序列标志以图形方式描绘。尽管序列标识可用于多种目的,但仅显示平均序列保守性,并且推断单个序列的保守性很困难。此处描述的个体信息 (R-i) 技术克服了这一限制。该方法首先根据比对序列的每个位置处的每个核苷酸或氨基酸的频率生成权重矩阵。然后将该矩阵应用于序列本身以确定每个单独序列的序列守恒性。该矩阵是唯一的,因为这些分配的平均值是总序列守恒,并且构建这样的矩阵只有一种方法。对于多核苷酸上的结合位点,权重矩阵具有将功能序列与其他序列区分开的自然截止值。 R-i 值是以信息位测量的绝对尺度,因此不同生物功能的守恒性可以相互比较。该矩阵可用于对序列进行排序、搜索新序列、将序列与其他定量数据(例如结合能或结合位点之间的距离)进行比较、区分突变与多态性、设计给定强度的序列以及检测数据库中的错误。 Ri 方法已用于识别以前未描述但经过实验验证的 DNA 结合位点。确定了大肠杆菌核糖体结合位点、细菌 Fis 结合位点以及人类供体和受体剪接点等的个体信息分布。这些分布清楚地表明共有序列非常不寻常,因此是描述自然发生的结合位点的一种糟糕方法。
Related genetic sequences having a common function can be described by Shannon's information measure and depicted graphically by a sequence logo. Though useful for many purposes, sequence logos only show the average sequence conservation, and inferring the conservation for individual sequences is difficult. This limitation is overcome by the individual information (R-i) technique described here. The method begins by generating a weight matrix from the frequencies of each nucleotide or amino acid at each position of the aligned sequences. This matrix is then applied to the sequences themselves to determine the sequence conservation of each individual sequence. The matrix is unique because the average of these assignments is the total sequence conservation, and there is only one way to construct such a matrix. For binding sites on polynucleotides, the weight matrix has a natural cut-off that distinguishes functional sequences from other sequences. R-i values are on an absolute scale measured in bits of information so the conservation of different biological functions can be compared with one another. The matrix can be used to rank-order the sequences, to search for new sequences, to compare sequences to other quantitative data such as binding energy or distance between binding sites, to distinguish mutations from polymorphisms, to design sequences of a given strength, and to detect errors in databases. The Ri method has been used to identify previously undescribed but experimentally verified DNA binding sites. The individual information distribution was determined for E, coli ribosome binding sites, bacterial Fis binding sites, and human donor and acceptor splice junctions, among others. The distributions demonstrate clearly that the consensus sequence is highly unusual, and hence is a poor method to describe naturally occurring binding sites.