Marginalized kernels for RNA sequence data analysis.

Marginalized kernels for RNA sequence data analysis.
复制标题

DOI:
10.11234/gi1990.13.112
复制
发表时间:
2002
期刊:
Genome informatics. International Conference on Genome Informatics
影响因子:
--
通讯作者:
Taishin Kin;K. Tsuda;K. Asai
Taishin Kin;K. Tsuda;K. Asai
中科院分区:
其他
文献类型:
--
作者:
Taishin Kin;K. Tsuda;K. Asai

文献摘要

相似文献

我们提出了新的内核,测量两个RNA序列的相似性,考虑到它们的二级结构。两种类型的内核。一种是已知二级结构的RNA序列,另一种是未知二级结构的RNA序列。后者采用随机上下文无关文法(SCFG)估计的二级结构。我们称后者为边缘化计数内核(MCK)。我们使用74组人类tRNA序列数据显示MCK的计算实验:(i)用于可视化tRNA相似性的核主成分分析(PCA),(ii)支持向量机(SVM)的监督分类。这两种类型的实验都显示出MCK的有希望的结果。
We present novel kernels that measure similarity of two RNA sequences, taking account of their secondary structures. Two types of kernels are presented. One is for RNA sequences with known secondary structures, the other for those without known secondary structures. The latter employs stochastic context-free grammar (SCFG) for estimating the secondary structure. We call the latter the marginalized count kernel (MCK). We show computational experiments for MCK using 74 sets of human tRNA sequence data: (i) kernel principal component analysis (PCA) for visualizing tRNA similarities, (ii) supervised classification with support vector machines (SVMs). Both types of experiment show promising results for MCKs.