Protein sequence comparison based on K-string dictionary

Protein sequence comparison based on K-string dictionary
复制标题

基于K字符串字典的蛋白质序列比对

DOI:
10.1016/j.gene.2013.07.092
复制
发表时间:
2013-10-01
期刊:
影响因子:
3.5
通讯作者:
Yau, Stephen S. -T.
Yau, Stephen S. -T.
中科院分区:
生物学3区
文献类型:
--
作者:
Yu, Chenglong;He, Rong L.;Yau, Stephen S. -T.

文献摘要

被引文献

相似文献

目前基于K串的蛋白质序列比较需要大量的计算机内存,因为蛋白质向量表示的维数随K呈指数增长。在本文中,我们提出了一个新的概念,“K串字典”,以解决这个高维问题。它允许我们使用一个低得多的维K-字符串为基础的频率或概率向量来表示蛋白质,从而显着降低其实现的计算机内存需求。在此基础上,我们利用奇异值分解对真实的蛋白质数据进行分析,改进的蛋白质向量表示方法可以得到精确的基因树。(C)2013爱思唯尔有限公司版权所有。
The current K-string-based protein sequence comparisons require large amounts of computer memory because the dimension of the protein vector representation grows exponentially with K. In this paper, we propose a novel concept, the "K-string dictionary", to solve this high-dimensional problem. It allows us to use a much lower dimensional K-string-based frequency or probability vector to represent a protein, and thus significantly reduce the computer memory requirements for their implementation. Furthermore, based on this new concept we use Singular Value Decomposition to analyze real protein datasets, and the improved protein vector representation allows us to obtain accurate gene trees. (C) 2013 Elsevier B.V. All rights reserved.