Towards the prediction of essential genes by integration of network topology, cellular localization and biological process information.

Towards the prediction of essential genes by integration of network topology, cellular localization and biological process information.
复制标题

通过整合网络拓扑,细胞定位和生物过程信息来预测必需基因。

DOI:
10.1186/1471-2105-10-290
复制
发表时间:
2009-09-16
期刊:
影响因子:
3
通讯作者:
Lemke N
Lemke N
中科院分区:
生物学4区
文献类型:
--
作者:
Acencio ML;Lemke N

文献摘要

参考文献

被引文献

相似文献

必需基因的鉴定对于理解细胞生命的最低要求和实际目的(例如药物设计)是重要的。然而,发现必需基因的实验技术是劳动密集型和耗时的。考虑到这些实验的限制,能够准确预测必需基因的计算方法将是非常有价值的。因此,我们在这里提出了一种基于机器学习的计算方法,依赖于网络拓扑特征,细胞定位和生物过程信息的预测的必需基因。我们构建了一个基于决策树的元分类器,并在具有个体和分组属性(网络拓扑特征,细胞区室和生物过程)的数据集上对其进行训练,以生成必需基因的各种预测因子。我们发现,具有更好性能的预测器是由具有集成属性的数据集生成的。使用具有所有属性的预测器,即,网络拓扑特征、细胞区室和生物学过程,我们获得了必需基因的最佳预测器,然后将其用于分类未知必需状态的酵母基因。最后,我们通过在具有所有网络拓扑特征、细胞定位和生物过程信息的数据集上训练J48算法来生成决策树,以发现本质性的细胞规则。我们发现,蛋白质物理相互作用的数量、蛋白质的核定位和转录调控因子的数量是决定基因重要性的最重要因素。我们能够证明,网络拓扑特征,细胞定位和生物过程信息是可靠的预测因子的必需基因。此外,通过构建基于这些数据的决策树,我们可以发现细胞规则的重要性。
The identification of essential genes is important for the understanding of the minimal requirements for cellular life and for practical purposes, such as drug design. However, the experimental techniques for essential genes discovery are labor-intensive and time-consuming. Considering these experimental constraints, a computational approach capable of accurately predicting essential genes would be of great value. We therefore present here a machine learning-based computational approach relying on network topological features, cellular localization and biological process information for prediction of essential genes. We constructed a decision tree-based meta-classifier and trained it on datasets with individual and grouped attributes-network topological features, cellular compartments and biological processes-to generate various predictors of essential genes. We showed that the predictors with better performances are those generated by datasets with integrated attributes. Using the predictor with all attributes, i.e., network topological features, cellular compartments and biological processes, we obtained the best predictor of essential genes that was then used to classify yeast genes with unknown essentiality status. Finally, we generated decision trees by training the J48 algorithm on datasets with all network topological features, cellular localization and biological process information to discover cellular rules for essentiality. We found that the number of protein physical interactions, the nuclear localization of proteins and the number of regulating transcription factors are the most important factors determining gene essentiality. We were able to demonstrate that network topological features, cellular localization and biological process information are reliable predictors of essential genes. Moreover, by constructing decision trees based on these data, we could discover cellular rules governing essentiality.
DOI: 10.1109/tkde.2005.50
发表时间: 2005-03-01
影响因子: 8.9
作者:
Huang, J;Ling, CX
通讯作者: Ling, CX
DOI: 10.1038/35075138
发表时间: 2001-05-03
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Mason, SP;Oltvai, ZN
通讯作者: Oltvai, ZN
DOI: 10.1093/nar/29.5.1144
发表时间: 2001-03-01
影响因子: 14.9
作者:
Daugeron, ML;Linder, P
通讯作者: Linder, P
DOI: 10.1002/j.1460-2075.1992.tb05099.x
发表时间: 1992-02-01
期刊: EMBO JOURNAL
影响因子: 11.4
作者:
GIRARD, JP;LEHTONEN, H;LAPEYRE, B
通讯作者: LAPEYRE, B
DOI: 10.1128/mcb.8.3.1067
发表时间: 1988-03-01
影响因子: 5.3
作者:
JACKSON, SP;LOSSKY, M;BEGGS, JD
通讯作者: BEGGS, JD