FermatS: A novel numerical representation for protein sequence comparison and DNA-binding protein identification

FermatS: A novel numerical representation for protein sequence comparison and DNA-binding protein identification
复制标题

FermatS:用于蛋白质序列比较和 DNA 结合蛋白鉴定的新型数值表示

DOI:
10.2174/1386207323999201117111738
复制
发表时间:
2021
期刊:
Combinatorial Chemistry & High Throughput Screening
影响因子:
--
通讯作者:
Xiaosheng Wang
Xiaosheng Wang
中科院分区:
其他
文献类型:
--
作者:
Yanping Zhang;Ya Gao;Jianwei Ni;Pengcheng Chen;Xiaosheng Wang

文献摘要

相似文献

目的:。基于蛋白质序列信息,采用了一种简单有效的方法。分析蛋白质序列相似性,预测dna结合蛋白。我们产生低复杂度的计算方法是绝对必要的。以准确地推断蛋白质的结构、功能和进化在迅速增长的数量。分子生物学数据可用目的:。产生新的计算算法进行分析和比较是很重要的。蛋白质序列随着分子生物学资料的不断增加而迅速增加。基于全局和局部位置表示的费马螺旋曲线和。分别给出了费马螺旋曲线的归一化转动惯量,以及组成。得到蛋白质序列的数值特征。它已被应用于分析9个ND5蛋白的相似性/差异性。分析结果与生物进化理论一致。此外,我们使用。Logistic回归与5倍交叉验证建立预测dna结合。在不平衡数据集中,该模型的F-measure值比dna abinder、DNA-prot、DNA-prot和gDNA-prot的F-measure值高0.0.0069 ~ 0.609,MCC值高0.293 ~ 0.898。这些结果表明,我们的方法,即费马,是有效的比较。蛋白质序列的识别与预测。
Aims:.Based on protein sequence information, a simple and effective method was used.to analyze protein sequence similarity and predict DNA-binding protein..Background:.It is absolutely necessary that we generate computational methods of low complexity.to accurate infer protein structure, function, and evolution in the rapidly growing number of.molecular biology data available..Objective:.It is important to generate novel computational algorithms for analyzing and comparing.protein sequences with the rapidly growing number of molecular biology data available..Method:.Based on global and local position representation with the curves of Fermat spiral and.normalized moments of inertia of the curve of Fermat spiral, respectively, moreover, composition.of 20 amino acids to get the numerical characteristics of protein sequences..Results:. It has been applied to analyze the similarity/dissimilarity of nine ND5 proteins, the.analysis results are consistent with the biological evolution theory. Furthermore, we employ the.Logistic regression with 5-fold cross-validation to establish the prediction of DNA-binding.proteins model, which outperformed the DNAbinder, iDNA-prot, DNA-prot and gDNA-prot by.0.0069-0.609 in terms of F-measure, 0.293-0.898 in terms of MCC in unbalanced dataset..Conclusion:.These results show that our method, namely FermatS, is effective to compare,.recognition and prediction the protein sequences.