GRAFENE: Graphlet-based alignment-free network approach integrates 3D structural and sequence (residue order) data to improve protein structural comparison.

GRAFENE: Graphlet-based alignment-free network approach integrates 3D structural and sequence (residue order) data to improve protein structural comparison.
复制标题

DOI:
10.1038/s41598-017-14411-y
复制
发表时间:
2017-11-02
期刊:
影响因子:
4.6
通讯作者:
Milenković T
Milenković T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Faisal FE;Newaz K;Chaney JL;Li J;Emrich SJ;Clark PL;Milenković T

文献摘要

参考文献

被引文献

相似文献

最初的蛋白质结构比较是基于序列的。由于序列中较远的氨基酸在 3 维 (3D) 结构中可以很接近,因此 3D 接触方法可以补充序列方法。传统的 3D 接触方法直接研究 3D 结构并且基于对齐。相反,3D 结构可以建模为蛋白质结构网络 (PSN)。然后,网络方法可以通过比较蛋白质的 PSN 来比较蛋白质。这些可以是基于对齐的或无对齐的。我们重点关注后者。现有的网络免对齐方法有缺点:1)它们依赖于网络拓扑的简单测量。 2) 它们对 PSN 大小不稳健。它们无法将 3) 多个 PSN 测量或 4) PSN 数据与序列数据集成,尽管这可以改善比较,因为不同的数据类型捕获蛋白质结构的互补方面。我们通过以下方式解决这个问题:1)通过新的网络免对齐方法利用完善的图基测量,2)引入归一化图基测量以消除 PSN 大小的偏差,3)允许集成多个 PSN 测量,以及 4)使用有序图基来组合互补的 PSN 数据和序列(特别是残差顺序)数据。我们比现有网络(无对齐和基于对齐)、3D 接触或序列方法更准确、更快速地比较合成网络和现实世界的 PSN。
Initial protein structural comparisons were sequence-based. Since amino acids that are distant in the sequence can be close in the 3-dimensional (3D) structure, 3D contact approaches can complement sequence approaches. Traditional 3D contact approaches study 3D structures directly and are alignment-based. Instead, 3D structures can be modeled as protein structure networks (PSNs). Then, network approaches can compare proteins by comparing their PSNs. These can be alignment-based or alignment-free. We focus on the latter. Existing network alignment-free approaches have drawbacks: 1) They rely on naive measures of network topology. 2) They are not robust to PSN size. They cannot integrate 3) multiple PSN measures or 4) PSN data with sequence data, although this could improve comparison because the different data types capture complementary aspects of the protein structure. We address this by: 1) exploiting well-established graphlet measures via a new network alignment-free approach, 2) introducing normalized graphlet measures to remove the bias of PSN size, 3) allowing for integrating multiple PSN measures, and 4) using ordered graphlets to combine the complementary PSN data and sequence (specifically, residue order) data. We compare synthetic networks and real-world PSNs more accurately and faster than existing network (alignment-free and alignment-based), 3D contact, or sequence approaches.
DOI: 10.1186/1471-2105-12-24
发表时间: 2011-01-19
期刊: BMC bioinformatics
影响因子: 3
作者:
Kuchaiev O;Stevanović A;Hayes W;Pržulj N
通讯作者: Pržulj N
DOI: 10.1089/cmb.2009.0196
发表时间: 2011-01-01
影响因子: 1.7
作者:
Andonov, Rumen;Malod-Dognin, Noel;Yanev, Nicola
通讯作者: Yanev, Nicola
DOI: 10.1110/ps.051479505
发表时间: 2005-08-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Kihara, D
通讯作者: Kihara, D
DOI: 10.1371/journal.pone.0003412
发表时间: 2008
期刊: PloS one
影响因子: 3.7
作者:
Clarke TF 4th;Clark PL
通讯作者: Clark PL
DOI: 10.1186/1477-5956-7-27
发表时间: 2009-08-09
期刊: Proteome science
影响因子: 2
作者:
Lee BJ;Shin MS;Oh YJ;Oh HS;Ryu KH
通讯作者: Ryu KH