Identifying statistical dependence in genomic sequences via mutual information estimates.

Identifying statistical dependence in genomic sequences via mutual information estimates.
复制标题

DOI:
10.1155/2007/14741
复制
发表时间:
2007
期刊:
EURASIP journal on bioinformatics & systems biology
影响因子:
--
通讯作者:
Szpankowski W
Szpankowski W
中科院分区:
其他
文献类型:
--
作者:
Aktulga HM;Kontoyiannis I;Lyznik LA;Szpankowski L;Grama AY;Szpankowski W

文献摘要

被引文献

相似文献

理解和量化生物体中信息的表示和数量的问题已成为生物学研究的核心部分,因为它们可能是根本性进步的关键。在本文中,我们展示了使用信息理论工具的任务,确定部分的生物分子(DNA或RNA)的统计相关。我们开发了一种精确可靠的方法,基于互信息的概念,用于查找和提取统计和结构依赖关系。定义了一个简单的阈值函数,并探讨了其在量化生物片段之间的依赖关系的显着性水平中的应用。这些工具用于两个特定的应用程序。首先,它们用于鉴定玉米zmSRp32基因的不同部分之间的相关性。在那里,我们发现zmSRp32的非翻译区和它的可变剪接外显子之间的显着依赖性。这一观察结果可能表明存在至今未知的选择性剪接机制或结构支架。第二,使用联邦调查局的联合DNA索引系统(CODIS)的数据,我们证明了我们的方法是特别适合的问题,发现短串联重复序列的重要性,在遗传分析的应用。
Questions of understanding and quantifying the representation and amount of information in organisms have become a central part of biological research, as they potentially hold the key to fundamental advances. In this paper, we demonstrate the use of information-theoretic tools for the task of identifying segments of biomolecules (DNA or RNA) that are statistically correlated. We develop a precise and reliable methodology, based on the notion of mutual information, for finding and extracting statistical as well as structural dependencies. A simple threshold function is defined, and its use in quantifying the level of significance of dependencies between biological segments is explored. These tools are used in two specific applications. First, they are used for the identification of correlations between different parts of the maize zmSRp32 gene. There, we find significant dependencies between the untranslated region in zmSRp32 and its alternatively spliced exons. This observation may indicate the presence of as-yet unknown alternative splicing mechanisms or structural scaffolds. Second, using data from the FBI's combined DNA index system (CODIS), we demonstrate that our approach is particularly well suited for the problem of discovering short tandem repeats—an application of importance in genetic profiling.