Network clustering coefficient approach to DNA sequence analysis

Network clustering coefficient approach to DNA sequence analysis
复制标题

DOI:
10.1016/j.chaos.2005.08.138
复制
发表时间:
2006-05
影响因子:
7.8
通讯作者:
G. Gerhardt;N. Lemke;G. Corso
G. Gerhardt;N. Lemke;G. Corso
中科院分区:
数学1区
文献类型:
--
作者:
G. Gerhardt;N. Lemke;G. Corso

文献摘要

被引文献

相似文献

在这项工作中,我们提出了一种基于图论概念的替代DNA序列分析工具。该方法通过三重网络研究生物体基因组的路径拓扑。在这个网络中,DNA序列中的三胞胎是顶点,如果两个顶点并列出现在基因组上,则它们是连接的。我们通过测量聚类系数来表征这种网络拓扑结构。我们针对两个主要偏差测试了我们的方法:鸟嘌呤-胞嘧啶(GC)含量和DNA序列的3-bp(碱基对)周期性。我们进行了测试,构建了可变GC含量的随机网络,并施加了3-bp的周期性。本文构建了一些生物体的测试群,并根据所构建的随机网络对该方法进行了研究。我们得出结论,聚类系数是一个有价值的工具,因为它提供的信息不平凡地包含在3-bp周期中,也不包含在可变的GC含量中。
In this work we propose an alternative DNA sequence analysis tool based on graph theoretical concepts. The methodology investigates the path topology of an organism genome through a triplet network. In this network, triplets in DNA sequence are vertices and two vertices are connected if they occur juxtaposed on the genome. We characterize this network topology by measuring the clustering coefficient. We test our methodology against two main bias: the guanine–cytosine (GC) content and 3-bp (base pairs) periodicity of DNA sequence. We perform the test constructing random networks with variable GC content and imposed 3-bp periodicity. A test group of some organisms is constructed and we investigate the methodology in the light of the constructed random networks. We conclude that the clustering coefficient is a valuable tool since it gives information that is not trivially contained in 3-bp periodicity neither in the variable GC content.