A machine learning approach for genome-wide prediction of morbid and druggable human genes based on systems-level data.

A machine learning approach for genome-wide prediction of morbid and druggable human genes based on systems-level data.
复制标题

DOI:
10.1186/1471-2164-11-s5-s9
复制
发表时间:
2010-12-22
期刊:
影响因子:
4.4
通讯作者:
Lemke N
Lemke N
中科院分区:
生物学2区
文献类型:
--
作者:
Costa PR;Acencio ML;Lemke N

文献摘要

被引文献

相似文献

在全基因组范围内识别病态基因(即那些突变导致遗传性人类疾病的基因)和可药物基因(即编码其被小分子调节引起表型效应的蛋白质的基因)都需要耗时和费力的实验方法。因此,一种能够在全基因组范围内准确预测此类基因的计算方法对于加快发现基因与疾病之间的因果关系以及确定基因产品的可药性将是非常宝贵的。在这篇文章中,我们提出了一种基于机器学习的计算方法来预测全基因组范围内的病态和可药物基因。为此,我们构建了一个基于决策树的元分类器,并在包含每个病态和可药物基因的网络拓扑特征、组织表达谱和亚细胞定位数据作为学习属性的数据集上对其进行训练。这个元分类器正确地恢复了65%的已知病态基因,精确度为66%,正确恢复了78%的已知可药物基因,精确度为75%。它被用来将发病率和可药性评分分配给未知的致病和可用药基因,我们显示了这些评分与文献数据之间的良好匹配。最后,我们通过在发病率和可药性数据集上训练J48算法来生成决策树,以发现发病率和可药性的细胞规则,并且在这些规则中,我们发现调节转录因子的数量和质膜定位分别是对发病率和可药性最重要的因素。我们能够证明,网络拓扑特征以及组织表达谱和亚细胞定位可以在全基因组范围内可靠地预测人类病态和可药物基因。此外,通过基于这些数据构建决策树,我们可以发现控制发病率和可药性的细胞规则。
The genome-wide identification of both morbid genes, i.e., those genes whose mutations cause hereditary human diseases, and druggable genes, i.e., genes coding for proteins whose modulation by small molecules elicits phenotypic effects, requires experimental approaches that are time-consuming and laborious. Thus, a computational approach which could accurately predict such genes on a genome-wide scale would be invaluable for accelerating the pace of discovery of causal relationships between genes and diseases as well as the determination of druggability of gene products. In this paper we propose a machine learning-based computational approach to predict morbid and druggable genes on a genome-wide scale. For this purpose, we constructed a decision tree-based meta-classifier and trained it on datasets containing, for each morbid and druggable gene, network topological features, tissue expression profile and subcellular localization data as learning attributes. This meta-classifier correctly recovered 65% of known morbid genes with a precision of 66% and correctly recovered 78% of known druggable genes with a precision of 75%. It was than used to assign morbidity and druggability scores to genes not known to be morbid and druggable and we showed a good match between these scores and literature data. Finally, we generated decision trees by training the J48 algorithm on the morbidity and druggability datasets to discover cellular rules for morbidity and druggability and, among the rules, we found that the number of regulating transcription factors and plasma membrane localization are the most important factors to morbidity and druggability, respectively. We were able to demonstrate that network topological features along with tissue expression profile and subcellular localization can reliably predict human morbid and druggable genes on a genome-wide scale. Moreover, by constructing decision trees based on these data, we could discover cellular rules governing morbidity and druggability.