Robust predictions of specialized metabolism genes through machine learning

Robust predictions of specialized metabolism genes through machine learning
复制标题

DOI:
10.1073/pnas.1817074116
复制
发表时间:
2019-02-05
影响因子:
11.1
通讯作者:
Shiu, Shin-Han
Shiu, Shin-Han
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Moore, Bethany M.;Wang, Peipei;Shiu, Shin-Han

文献摘要

被引文献

相似文献

植物特化代谢(SM)酶产生具有重要生态学、进化和生物技术意义的谱系特异性代谢物。以拟南芥(Arabidopsis thaliana)为模型,通过对SM和GM(general metabolism,一般称为初级代谢)基因的复制模式、序列保守性、转录、蛋白质结构域含量和基因网络特性等特征的详细研究,确定了SM和GM基因的区别特征。多套基准基因的分析表明,SM基因往往是串联重复,共表达与他们的旁系同源,狭义地表达在较低的水平,保守性较低,以及在基因网络中连接不太好相对于GM基因。虽然SM和GM基因之间的这些功能的值显着不同,任何单一的功能是无效的,从GM基因预测SM。使用机器学习方法整合所有特征,建立了真阳性率为87%,真阴性率为71%的预测模型。此外,86%的已知SM基因没有被用于创建机器学习模型。我们还表明,该模型可以进一步改善,当我们区分SM,GM和交界处的基因负责SM和GM途径共享的反应,表明拓扑考虑可能会进一步改善SM预测模型。应用预测模型识别出1,220 A。thaliana基因与以前未知的功能,每个分配一个称为SM得分的置信度措施,提供了一个全球估计SM基因在植物基因组中的含量。
Plant specialized metabolism (SM) enzymes produce lineage-specific metabolites with important ecological, evolutionary, and biotechnological implications. Using Arabidopsis thaliana as a model, we identified distinguishing characteristics of SM and GM (general metabolism, traditionally referred to as primary metabolism) genes through a detailed study of features including duplication pattern, sequence conservation, transcription, protein domain content, and gene network properties. Analysis of multiple sets of benchmark genes revealed that SM genes tend to be tandemly duplicated, coexpressed with their paralogs, narrowly expressed at lower levels, less conserved, and less well connected in gene networks relative to GM genes. Although the values of each of these features significantly differed between SM and GM genes, any single feature was ineffective at predicting SM from GM genes. Using machine learning methods to integrate all features, a prediction model was established with a true positive rate of 87% and a true negative rate of 71%. In addition, 86% of known SM genes not used to create the machine learning model were predicted. We also demonstrated that the model could be further improved when we distinguished between SM, GM, and junction genes responsible for reactions shared by SM and GM pathways, indicating that topological considerations may further improve the SM prediction model. Application of the prediction model led to the identification of 1,220 A. thaliana geneswith previously unknown functions, each assigned a confidence measure called an SM score, providing a global estimate of SM gene content in a plant genome.