Splice site identification using probabilistic parameters and SVM classification.

Splice site identification using probabilistic parameters and SVM classification.
复制标题

DOI:
10.1186/1471-2105-7-s5-s15
复制
发表时间:
2006-12-18
期刊:
影响因子:
3
通讯作者:
Li J
Li J
中科院分区:
生物学4区
文献类型:
--
作者:
Baten AK;Chang BC;Halgamuge SK;Li J

文献摘要

被引文献

相似文献

DNA 测序技术的最新进展和自动化创造了大量的 DNA 序列数据。序列数据的不断增长需要更好、更高效的分析方法。在这些新积累的数据中识别基因是生物信息学的一个重要问题,它需要预测完整的基因结构。 DNA 序列中剪接位点的准确识别在真核生物基因结构预测中发挥着核心作用之一。剪接位点的有效检测需要了解剪接位点周围区域中核苷酸的特征、依赖性和关系。高阶马尔可夫模型通常被认为是建模高阶依赖关系的有用技术。然而,它们的实现需要估计大量参数,这在计算上是昂贵的。所提出的剪接位点检测方法由两个阶段组成:第一阶段使用一阶马尔可夫模型(MM1),第二阶段使用具有多项式核的支持向量机(SVM)。 MM1 充当 SVM 的预处理步骤,并将 DNA 序列作为其输入。它根据剪接位点区域周围的概率参数对核苷酸的组成特征和依赖性进行建模。然后将概率参数输入支持向量机,将它们非线性组合以预测剪接位点。当所提出的 MM1-SVM 模型与其他现有的标准剪接位点检测方法进行比较时,它在所有情况下都显示出优越的性能。我们提出了一种有效的支持向量机预处理方案,并将其应用于剪接位点的识别。这是一种简单而有效的剪接位点检测方法,与其他一些更复杂的方法相比,它表现出更好的分类精度和计算速度。
Recent advances and automation in DNA sequencing technology has created a vast amount of DNA sequence data. This increasing growth of sequence data demands better and efficient analysis methods. Identifying genes in this newly accumulated data is an important issue in bioinformatics, and it requires the prediction of the complete gene structure. Accurate identification of splice sites in DNA sequences plays one of the central roles of gene structural prediction in eukaryotes. Effective detection of splice sites requires the knowledge of characteristics, dependencies, and relationship of nucleotides in the splice site surrounding region. A higher-order Markov model is generally regarded as a useful technique for modeling higher-order dependencies. However, their implementation requires estimating a large number of parameters, which is computationally expensive. The proposed method for splice site detection consists of two stages: a first order Markov model (MM1) is used in the first stage and a support vector machine (SVM) with polynomial kernel is used in the second stage. The MM1 serves as a pre-processing step for the SVM and takes DNA sequences as its input. It models the compositional features and dependencies of nucleotides in terms of probabilistic parameters around splice site regions. The probabilistic parameters are then fed into the SVM, which combines them nonlinearly to predict splice sites. When the proposed MM1-SVM model is compared with other existing standard splice site detection methods, it shows a superior performance in all the cases. We proposed an effective pre-processing scheme for the SVM and applied it for the identification of splice sites. This is a simple yet effective splice site detection method, which shows a better classification accuracy and computational speed than some other more complex methods.