Using hidden Markov models to investigate G-quadruplex motifs in genomic sequences.

Using hidden Markov models to investigate G-quadruplex motifs in genomic sequences.
复制标题

DOI:
10.1186/1471-2164-15-s9-s15
复制
发表时间:
2014
期刊:
影响因子:
4.4
通讯作者:
Kato Y
Kato Y
中科院分区:
生物学2区
文献类型:
--
作者:
Yano M;Kato Y

文献摘要

相似文献

G-四链体是在富含鸟嘌呤的核苷酸序列中形成的四链结构。到目前为止,DNA G-四链体的几种功能作用已经被研究,其中它们在DNA复制和转录过程中的假定功能作用已经被提出。G-四链体形成的必要条件是存在四个称为G-链的串联鸟嘌呤区域和连接G-链的三个称为环的核苷酸序列。检测给定基因组序列中潜在的G-四链体区域的一种简单的计算方法是使用正则表达式进行模式匹配。尽管通过基于常规表达的方法可以在大多数基因组中发现许多推定的G-四链体基序,但是这些序列中的大多数不太可能形成G-四链体,因为它们与典型的双螺旋结构相比是不稳定的。在这里,我们提出了精心设计的计算模型,用于表示DNA G-四链体基序使用隐马尔可夫模型(HALGORY)。使用的阻碍使我们能够评估G-四链体图案定量的概率措施。此外,可以通过使用实验验证的数据来训练Hestival的参数。计算实验在区分正和负的G-四链体序列,以及减少推定的G-四链体在人类基因组中进行,表明基于HMM的模型可以辨别真正的G-四链体结构以及其中之一有可能减少假阳性G-四链体预测现有的正则表达式为基础的方法。此外,我们的研究结果表明,我们的模型之一可以专门检测G-四链体序列,其功能作用预计将参与DNA转录。基于HMM的方法沿着传统的模式匹配方法可以有助于减少对给定的一组感兴趣的潜在G-四链体进行功能分析的昂贵且费力的湿实验室实验。C++和Perl程序可以在http://tcs.cira.kyoto-u.ac.jp/~ykato/program/g4hmm/上找到。
G-quadruplexes are four-stranded structures formed in guanine-rich nucleotide sequences. Several functional roles of DNA G-quadruplexes have so far been investigated, where their putative functional roles during DNA replication and transcription have been suggested. A necessary condition for G-quadruplex formation is the presence of four regions of tandem guanines called G-runs and three nucleotide subsequences called loops that connect G-runs. A simple computational way to detect potential G-quadruplex regions in a given genomic sequence is pattern matching with regular expression. Although many putative G-quadruplex motifs can be found in most genomes by the regular expression-based approach, the majority of these sequences are unlikely to form G-quadruplexes because they are unstable as compared with canonical double helix structures. Here we present elaborate computational models for representing DNA G-quadruplex motifs using hidden Markov models (HMMs). Use of HMMs enables us to evaluate G-quadruplex motifs quantitatively by a probabilistic measure. In addition, the parameters of HMMs can be trained by using experimentally verified data. Computational experiments in discriminating between positive and negative G-quadruplex sequences as well as reducing putative G-quadruplexes in the human genome were carried out, indicating that HMM-based models can discern bona fide G-quadruplex structures well and one of them has the possibility of reducing false positive G-quadruplexes predicted by existing regular expression-based methods. Furthermore, our results show that one of our models can be specialized to detect G-quadruplex sequences whose functional roles are expected to be involved in DNA transcription. The HMM-based method along with the conventional pattern matching approach can contribute to reducing costly and laborious wet-lab experiments to perform functional analysis on a given set of potential G-quadruplexes of interest. The C++ and Perl programs are available at http://tcs.cira.kyoto-u.ac.jp/~ykato/program/g4hmm/.