Optimising the use of gene expression data to predict plant metabolic pathway memberships

Optimising the use of gene expression data to predict plant metabolic pathway memberships
复制标题

优化基因表达数据的使用来预测植物代谢途径成员资格

DOI:
10.1111/nph.17355
复制
发表时间:
2021
期刊:
影响因子:
9.4
通讯作者:
Shiu, Shin‐Han
Shiu, Shin‐Han
中科院分区:
生物学1区
文献类型:
--
作者:
Wang, Peipei;Moore, Bethany M.;Uygun, Sahra;Lehti‐Shiu, Melissa D.;Barry, Cornelius S.;Shiu, Shin‐Han

文献摘要

相似文献

多种途径的植物代谢物对植物的生存、人类的营养和医学都具有重要意义。大多数植物酶基因的通路成员是未知的。虽然共表达对于基因通路的分配是有用的,但表达相关性可能只存在于特定的时空和条件背景下。利用bbb600番茄(Solanum lycopersicum)表达数据组合,探索了85条通路中预测成员关系的三种策略。对不同路径的最优预测需要不同的数据组合来指示路径函数。幼稚的预测(即用表达最相似的基因识别途径)容易出错。在52条路径中,无监督学习比有监督学习表现更好,可能是由于训练数据的可用性有限。使用基因-通路表达相似性导致预测模型优于单纯基于表达水平的预测模型。使用36个实验验证的基因,路径最佳模型预测准确率为58.3%,明显优于没有实验证据的注释基因预测(37.0%)或随机猜测(1.2%),表明数据质量的重要性。我们的研究强调需要广泛探索基于表达的特征和预测策略,以最大限度地提高代谢途径成员分配的准确性。这里概述的预测框架可以应用于其他物种,并作为未来比较的基线模型。
Plant metabolites from diverse pathways are important for plant survival, human nutrition and medicine. The pathway memberships of most plant enzyme genes are unknown. While co‐expression is useful for assigning genes to pathways, expression correlation may exist only under specific spatiotemporal and conditional contexts.Utilising > 600 tomato (Solanum lycopersicum) expression data combinations, three strategies for predicting memberships in 85 pathways were explored.Optimal predictions for different pathways require distinct data combinations indicative of pathway functions. Naive prediction (i.e. identifying pathways with the most similarly expressed genes) is error prone. In 52 pathways, unsupervised learning performed better than supervised approaches, possibly due to limited training data availability. Using gene‐to‐pathway expression similarities led to prediction models that outperformed those based simply on expression levels. Using 36 experimental validated genes, the pathway‐best model prediction accuracy is 58.3%, significantly better compared with that for predicting annotated genes without experimental evidence (37.0%) or random guess (1.2%), demonstrating the importance of data quality.Our study highlights the need to extensively explore expression‐based features and prediction strategies to maximise the accuracy of metabolic pathway membership assignment. The prediction framework outlined here can be applied to other species and serves as a baseline model for future comparisons.