Computational Identification of Essential Genes in Prokaryotes and Eukaryotes

Computational Identification of Essential Genes in Prokaryotes and Eukaryotes
复制标题

DOI:
10.1007/978-3-319-94806-5_13
复制
发表时间:
2017-02
期刊:
--
影响因子:
--
通讯作者:
Dawit Nigatu;W. Henkel
Dawit Nigatu;W. Henkel
中科院分区:
其他
文献类型:
--
作者:
Dawit Nigatu;W. Henkel

文献摘要

相似文献

提出了几种用于鉴定必需基因(EG)的计算方法。基于机器学习的方法使用源自基因序列、基因表达数据、网络拓扑、同源性和域信息的特征。除了基于序列的特征外,其他特征都需要额外的实验数据,而这些数据对于尚未充分研究和新测序的生物体来说是无法获得的。因此,在这里,我们提出了基于序列的 EG 识别。我们对 15 种细菌、1 种古细菌和 4 种真核生物进行了基因必要性预测。信息论量,例如互信息、条件互信息、熵、Kullback-Leibler 散度和马尔可夫模型,被用作特征。此外,为了提高预测性能,还包括与终止密码子使用、长度和 GC 内容相关的其他易于访问的基于序列的特征。对于分类,使用随机森林算法。通过采用组织内和跨组织预测来广泛评估所提出方法的性能。所获得的结果优于大多数先前发表的仅依赖于序列信息的 EG 预测器,并且与使用源自网络拓扑、同源性和基因表达数据的附加特征的结果相当。
Several computational methods were proposed for the identification of essential genes (EGs). The machine learning based methods use features derived from the genetic sequences, gene-expression data, network topology, homology, and domain information. Except for the sequence-based features, the others require additional experimental data which is unavailable for under-studied and newly sequenced organisms. Hence, here, we propose a sequence-based identification of EGs. We performed gene essentiality predictions considering 15 bacteria, 1 archeaon, and 4 eukaryotes. Information-theoretic quantities, such as mutual information, conditional mutual information, entropy, Kullback-Leibler divergence, and Markov models, were used as features. In addition, with the hope of improving the prediction performance, other easily accessible sequence-based features related to stop codon usage, length, and GC content were included. For classification, the Random Forest algorithm was used. The performance of the proposed method is extensively evaluated by employing both intra- and cross-organism predictions. The obtained results were better than most of the previously published EG predictors which rely only on sequence information and comparable to those using additional features derived from network topology, homology, and gene-expression data.