Profile-based string kernels for remote homology detection and motif extraction

Profile-based string kernels for remote homology detection and motif extraction
复制标题

DOI:
10.1142/s021972000500120x
复制
发表时间:
2005-06-01
影响因子:
1
通讯作者:
Leslie, Christina
Leslie, Christina
中科院分区:
生物学4区
文献类型:
--
作者:
Kuang, Rui;Ie, Eugene;Leslie, Christina

文献摘要

被引文献

相似文献

我们介绍了新的配置文件为基础的字符串内核与支持向量机(SVM)的蛋白质分类和远程同源性检测的问题。这些核使用概率分布,例如由PSI-BLAST算法产生的概率分布,来定义沿着沿着蛋白质序列的位置依赖性突变邻域,以用于数据中k长度序列(“k-mers”)的不精确匹配。通过使用有效的数据结构,一旦已经获得轮廓,核就可以快速计算。例如,运行PSI-BLAST以构建配置文件所需的时间明显长于内核计算时间和SVM训练时间。我们提出了远程同源性检测实验的SCOP数据库的基础上,我们表明,基于配置文件的字符串内核与SVM分类器的性能大大优于所有最近提出的监督SVM方法。我们进一步研究如何将预测的二级结构信息纳入配置文件内核,以获得一个小的,但显着的性能改善。我们还展示了如何使用学习的SVM分类器来提取“区分序列基序”-原始配置文件的短区域,这些区域几乎占SVM分类得分的所有权重,并表明这些区分基序对应于蛋白质数据中有意义的结构特征。PSI-BLAST图谱的使用可以被视为半监督学习技术,因为PSI-BLAST利用来自大型序列数据库的未标记数据来构建更多信息图谱。最近提出的“聚类核”给出了一般的半监督方法,用于提高SVM蛋白质分类性能。我们表明,我们的配置文件内核的结果也优于集群内核,同时提供更好的可扩展性,大型数据集。
We introduce novel profile-based string kernels for use with support vector machines (SVMs) for the problems of protein classification and remote homology detection. These kernels use probabilistic profiles, such as those produced by the PSI-BLAST algorithm, to define position-dependent mutation neighborhoods along protein sequences for inexact matching of k-length subsequences ("k-mers") in the data. By use of an efficient data structure, the kernels are fast to compute once the profiles have been obtained. For example, the time needed to run PSI-BLAST in order to build the profiles is significantly longer than both the kernel computation time and the SVM training time. We present remote homology detection experiments based on the SCOP database where we show that profile-based string kernels used with SVM classifiers strongly outperform all recently presented supervised SVM methods. We further examine how to incorporate predicted secondary structure information into the profile kernel to obtain a small but significant performance improvement. We also show how we can use the learned SVM classifier to extract "discriminative sequence motifs" - short regions of the original profile that contribute almost all the weight of the SVM classification score and show that these discriminative motifs correspond to meaningful structural features in the protein data. The use of PSI-BLAST profiles can be seen as a semi-supervised learning technique, since PSI-BLAST leverages unlabeled data from a large sequence database to build more informative profiles. Recently presented "cluster kernels" give general semi-supervised methods for improving SVM protein classification performance. We show that our profile kernel results also outperform cluster kernels while providing much better scalability to large datasets.