A multi-approaches-guided genetic algorithm with application to operon prediction

A multi-approaches-guided genetic algorithm with application to operon prediction
复制标题

一种应用于操纵子预测的多方法引导遗传算法

DOI:
10.1016/j.artmed.2007.07.010
复制
发表时间:
2007-10-01
影响因子:
7.5
通讯作者:
Liang, Yanchun
Liang, Yanchun
中科院分区:
工程技术1区
文献类型:
--
作者:
Wang, Shuqin;Wang, Yan;Liang, Yanchun

文献摘要

被引文献

相似文献

目的:操纵子的预测是在全基因组水平上重建调控网络的关键。多基因组特征已被用于预测操纵子。然而,文献中通常仅使用单一方法来处理多个基因组特征。本文的目的是通过对不同基因组特征采用不同的预处理方法,以发挥其独特的特征,从而发展一种操纵子预测的组合方法。我们利用不同的方法,基因间距离,集群的orthopathic组(COG)基因功能,代谢途径和微阵列表达数据。提出了一种新的局部熵最小化方法来划分基因间距离。我们的程序可以用于其他新测序的基因组转移的知识,已从大肠杆菌的数据。我们计算COG基因函数的对数似然和微阵列表达数据的Pearson相关系数。结果:在E。coli K12基因组、枯草芽孢杆菌基因组和铜绿假单胞菌PAO 1基因组。对这三个基因组的预测准确率分别为85.9987%、88.296%和81.2384%。结论:模拟实验结果表明,在遗传算法中,采用多种方法对基因组数据进行预处理,保证了不同生物特征的有效利用。实验结果也表明,该方法适用于预测原核生物中的操纵子。(C)2007 Elsevier B.V.保留所有权利。
Objective: The prediction of operons is critical to the reconstruction of regulatory networks at the whole genome level. Multiple genome features have been used for predicting operons. However, multiple genome features are usually dealt with using only single method in the literatures. The aim of this paper is to develop a combined method for operon prediction by using different methods to preprocess different genome features in order for exerting their unique characteristics.Methods: A novel multi-approach-guided genetic algorithm for operon prediction is presented. We exploit different methods for intergenic distance, cluster of orthologous groups (COG) gene functions, metabolic pathway and microarray expression data. A novel local-entropy-minimization method is proposed to partition intergenic distance. Our program can be used for other newly sequenced genomes by transferring the knowledge that has been obtained from Escherichia coli data. We calculate the log-likelihood for COG gene functions and Pearson correlation coefficient for microarray expression data. The genetic algorithm is used for integrating the four types of data.Results: The proposed method is examined on E. coli K12 genome, Bacillus subtilis genome, and Pseudomonas aeruginosa PAO1 genome. The accuracies of prediction for these three genomes are 85.9987%, 88.296%, and 81.2384%, respectively.Conclusion: Simulated experimental results demonstrate that in the genetic algorithm the preprocessing for genome data using multiple approaches ensures the effective utilization of different biological characteristics. Experimental results also show that the proposed method is applicable for predicting operons in prokaryote. (C) 2007 Elsevier B.V. All rights reserved.