A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysis.

A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysis.
复制标题

DOI:
10.1186/1471-2105-9-510
复制
发表时间:
2008-12-01
期刊:
影响因子:
3
通讯作者:
Wang X
Wang X
中科院分区:
生物学4区
文献类型:
--
作者:
Liu B;Wang X;Lin L;Dong Q;Wang X

文献摘要

参考文献

被引文献

相似文献

蛋白质远程同源性检测和折叠识别是生物信息学研究的核心问题。目前,基于支持向量机(SVM)的判别方法是解决这些问题最有效、最准确的方法。提高基于支持向量机的方法性能的关键步骤是找到合适的蛋白质序列表示。本文提出了一种新的蛋白质构建模块Top-n-grams,它包含了从蛋白质序列频率谱中提取的进化信息。从PSI-BLAST输出的多个序列比对中计算蛋白质序列频率谱,并将其转换为top -n-gram。根据Top-n-gram的出现次数将蛋白质序列转化为固定维的特征向量。利用支持向量机对训练向量进行评估,训练分类器对测试蛋白序列进行分类。研究表明,将Top-n-grams和潜在语义分析(LSA)相结合,可以提高远程同源检测和折叠识别的预测性能,这是一种有效的自然语言处理特征提取技术。在超家族和折叠基准测试中,结合Top-n-grams和LSA的方法的结果明显优于相关方法。基于Top-n-grams的方法明显优于基于N-grams、pattern、motif和binary profile等许多其他构建块的方法。因此,Top-n-gram是一种很好的蛋白质序列构建模块,可广泛应用于计算生物学的许多任务,如序列比对、结构域边界预测、知识势的指定和蛋白质结合位点的预测。
Protein remote homology detection and fold recognition are central problems in bioinformatics. Currently, discriminative methods based on support vector machine (SVM) are the most effective and accurate methods for solving these problems. A key step to improve the performance of the SVM-based methods is to find a suitable representation of protein sequences. In this paper, a novel building block of proteins called Top-n-grams is presented, which contains the evolutionary information extracted from the protein sequence frequency profiles. The protein sequence frequency profiles are calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into Top-n-grams. The protein sequences are transformed into fixed-dimension feature vectors by the occurrence times of each Top-n-gram. The training vectors are evaluated by SVM to train classifiers which are then used to classify the test protein sequences. We demonstrate that the prediction performance of remote homology detection and fold recognition can be improved by combining Top-n-grams and latent semantic analysis (LSA), which is an efficient feature extraction technique from natural language processing. When tested on superfamily and fold benchmarks, the method combining Top-n-grams and LSA gives significantly better results compared to related methods. The method based on Top-n-grams significantly outperforms the methods based on many other building blocks including N-grams, patterns, motifs and binary profiles. Therefore, Top-n-gram is a good building block of the protein sequences and can be widely used in many tasks of the computational biology, such as the sequence alignment, the prediction of domain boundary, the designation of knowledge-based potentials and the prediction of protein binding sites.
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
DOI: 10.1142/s021972000500120x
发表时间: 2005-06-01
影响因子: 1
作者:
Kuang, Rui;Ie, Eugene;Leslie, Christina
通讯作者: Leslie, Christina
DOI: 10.1089/106652703322756113
发表时间: 2003-01-01
影响因子: 1.7
作者:
Liao, L;Noble, WS
通讯作者: Noble, WS
DOI: 10.1093/bioinformatics/bti801
发表时间: 2006-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dong, QW;Wang, XL;Lin, L
通讯作者: Lin, L
Windows .NET网络分布式基本本地对齐搜索工具包(W.ND-Blast)。
DOI: 10.1186/1471-2105-6-93
发表时间: 2005-04-08
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Dowd, SE;Zaragoza, J;Rodriguez, JR;Oliver, MJ;Payton, PR
通讯作者: Payton, PR