Protein sequence comparison based on K-string dictionary
Protein sequence comparison based on K-string dictionary
复制标题
基于K字符串字典的蛋白质序列比对
DOI:
10.1016/j.gene.2013.07.092
复制
发表时间:
2013-10-01
期刊:
影响因子:
3.5
通讯作者:
Yau, Stephen S. -T.
中科院分区:
文献类型:
--
作者:
Yu, Chenglong;He, Rong L.;Yau, Stephen S. -T.
The current K-string-based protein sequence comparisons require large amounts of computer memory because the dimension of the protein vector representation grows exponentially with K. In this paper, we propose a novel concept, the "K-string dictionary", to solve this high-dimensional problem. It allows us to use a much lower dimensional K-string-based frequency or probability vector to represent a protein, and thus significantly reduce the computer memory requirements for their implementation. Furthermore, based on this new concept we use Singular Value Decomposition to analyze real protein datasets, and the improved protein vector representation allows us to obtain accurate gene trees. (C) 2013 Elsevier B.V. All rights reserved.