Normalized Feature Vectors: A Novel Alignment-Free Sequence Comparison Method Based on the Numbers of Adjacent Amino Acids

Normalized Feature Vectors: A Novel Alignment-Free Sequence Comparison Method Based on the Numbers of Adjacent Amino Acids
复制标题

DOI:
10.1109/tcbb.2013.10
复制
发表时间:
2013-03-01
影响因子:
4.5
通讯作者:
Yu, Hong-Jie
Yu, Hong-Jie
中科院分区:
工程技术3区
文献类型:
--
作者:
Huang, De-Shuang;Yu, Hong-Jie

文献摘要

被引文献

相似文献

基于各种相邻氨基酸(AAA),我们将每个蛋白质一级序列映射到400 ×(L - 1)矩阵M中。此外,我们进一步推导出一个归一化的400元组的数学描述符D,它是通过矩阵的奇异值分解(SVD)从初级蛋白质序列中提取的。所获得的400维归一化特征向量(NFV)进一步促进了我们对蛋白质序列的定量分析。利用一级蛋白质序列的归一化表示法,对9个物种的ND 5序列和24种脊椎动物的转铁蛋白序列进行了相似性分析。我们也比较了本研究的结果与其他相关工作。这两个实验表明,我们提出的NFV-AAA方法在序列相似性分析领域表现良好。
Based on all kinds of adjacent amino acids (AAA), we map each protein primary sequence into a 400 by (L - 1) matrix M. In addition, we further derive a normalized 400-tuple mathematical descriptors D, which is extracted from the primary protein sequences via singular values decomposition (SVD) of the matrix. The obtained 400-D normalized feature vectors (NFVs) further facilitate our quantitative analysis of protein sequences. Using the normalized representation of the primary protein sequences, we analyze the similarity for different sequences upon two data sets: 1) ND5 sequences from nine species and 2) transferrin sequences of 24 vertebrates. We also compared the results in this study with those from other related works. These two experiments illustrate that our proposed NFV-AAA approach does perform well in the field of similarity analysis of sequence.